An Uncertainty Risk Evaluation Tool for Wellbore Leakage Prediction for Plug and Abandonement (P&A)
Bibliographic record
Abstract
Abstract Oil and gas wells leakage is a major concern due to the associated risks. Potential issues include habitat fragmentation, soil erosion, groundwater contamination, and greenhouse gas emissions released into the atmosphere. An estimated 2 million abandoned oil and gas wells are believed to be leakage. Proper Plug and Abandonment (P&A) operations are required to ensure these wells are correctly disposed of from their useful operational life. This study aims to build an uncertainty evaluation tool to statistically classify the risk of a well from leaking based on their well information (age, location, depth, completion interval, casings, and cement). Data consists of leakage reports and available well data reports from Alberta Energy Regulator (AER) in Canada. Multiple preprocessing techniques, including balancing the data, encoding, and standardization, were implemented before training. Multiple models that included Naïve Bayes (NB), Support Vector Machine (SVM), Decision Trees (DT), Random Forest (RF), and K-Nearest Neighbors (KNN) were compared to select the best-performing for optimization. RF outperformed the other models and was tuned using hyperparameter optimization and cross-validation. The final model's average accuracy was 77.1% across all folds. Multiple evaluation metrics, including Accuracy, Confusion Matrix, Precision, Recall, and Area Under the ROC Curve (AUC), were used to assess the model and each class against the rest. Feature importance showed an even distribution across the different features used. The model presented in the study aimed to classify wells and label the leakage risk based on the well information associated with its components. This risk evaluation tool could help reduce gas emissions by 28.2% based on the results obtained. This tool can classify the wells to speed the selection process and prioritize wells with higher leakage risk to perform P&A operations and minimize emissions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".