Objective Probability of Detection Functions for High-Resolution Detection of Point-Source Methane Emissions from Space 
Bibliographic record
Abstract
Reliable detection and localization of methane point sources is critical to rapidly mitigating methane emissions and avoiding the worst-case projected scenarios of climate change. With increasingly widespread coverage and short revisit periods, satellite-based detection of methane point-sources is well-suited to identifying the largest emitters for mitigation. To fully understand how satellite-based detection fits within multi-scale monitoring, reporting, and verification (MRV) frameworks, however, it is crucial to robustly characterize the probability of detection (POD) of active and planned instruments. This is especially true in the upstream oil and gas sector, where measurements show highly right-skewed emissions size distributions that raise doubt on the portion of sector-wide emissions detectable from space. While there is no shortage of statements regarding vaguely defined limits of detection (LODs) or PODs within the literature and press releases, these are often opaquely derived and, crucially, are not placed in the context of false positive rates. Moreover, LODs based on single-blind controlled release testing are inherently problematic as prior knowledge of controlled releases precludes objective interpretation of detectability. With ongoing efforts to improve detection sensitivities and quantification accuracies, a standardized framework for objective and transparent PODs is needed to ensure satellite-based instruments are leveraged to their full capacity. Here, we present an innovative framework to objectively quantify detection probabilities for satellite observations of methane point sources. This generalized framework is independent of satellite resolution and precision in addition to the underlying – and potentially proprietary – method for plume segmentation. Based in frequentist hypothesis testing, the new framework provides a POD function for satellite-specific characteristics, given a permissible false positive rate. This latter point is critical when considering satellite techniques for use in MRV as well as leak detection and repair (LDAR) programs where nuisance detections (i.e., false positives) can render satellite approaches cost prohibitive. We apply this new framework to methane-detecting satellite instruments in two scenarios: 1) idealized conditions of spatially uncorrelated image noise and 2) realistic conditions with spatially correlated noise derived from publicly available image data. These analyses identify best-case PODs as a function of imagery resolution and precision and show the detrimental effect of spatially correlated noise on automated plume segmentation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.020 | 0.056 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.003 | 0.004 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.010 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".