Automatic error correction: Improving annotation quality for model optimization in oil-exploration related land disturbances mapping
Bibliographic record
Abstract
The manual extraction of land disturbances associated with oil exploration, which normally includes resource roads, mining facilities, and well pads, presents significant challenges in terms of cost and time. Accurate monitoring and mapping of land disturbances resulting from oil exploration plays a crucial role in conducting comprehensive environmental assessments and facilitating effective land reclamation initiatives. However, prevailing deep learning methodologies in the realm of oil and gas exploration primarily focus on oil spill detection, neglecting the critical aspect of land disturbances resulting from oil exploration, thus overlooking the impact on land. Furthermore, given that the well sites are scattered and relatively diminutive compared to other land covers, their detection poses substantial difficulties. This paper proposes an automatic error-correcting (AEC) algorithm to address deficiencies in ground truth data quality. This AEC method was integrated into the deep-learning framework for land disturbance extraction, specifically tailored for land disturbances analysis associated with oil exploration. The efficacy of our method was validated on a dataset collected in Alberta covering an area of oil sand mining sites. The application of the AEC algorithm significantly enhanced the accuracy of land disturbance analysis, thereby contributing to a more effective hydrocarbon exploration impact analysis and facilitating the timely planning by the Alberta government. The results demonstrate notable improvements in both average pixel accuracy (AA) and mean intersection over union (mIoU), ranging from 8.3% to 15.4% and 0.5% to 5.8%, respectively. These enhancements, which have profound implications for the precision of land disturbance detection, prove that the proposed AEC algorithm can serve a dual purpose: correcting errors in the dataset and efficiently detecting land disturbance features in the oil exploration area.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".