Accounting for Non-Detects: Application to Satellite Ammonia Observations
Bibliographic record
Abstract
Presented is a methodology to explicitly identify and account for cloud-free satellite measurements below a sensor’s measurement detection level. These low signals can often be found in satellite observations of minor atmospheric species with weak spectral signals (e.g., ammonia (NH3)). Not accounting for these non-detects can high-bias averaged measurements in locations that exhibit conditions below the detection limit of the sensor. The approach taken here is to utilize the information content from the satellite signal to explicitly identify non-detects and then account for them with a consistent approach. The methodology is applied to the CrIS Fast Physical Retrieval (CFPR) ammonia product and results in a more realistic averaged dataset under conditions where there are a significant number of non-detects. These results show that in larger emission source regions (i.e., surface values > 7.5 ppbv) the non-detects occur less than 5% of the time and have a relatively small impact (decreases by less than 5%) on the gridded averaged values (e.g., annual ammonia source regions). However, in regions that have low ammonia concentration amounts (i.e., surface values < 1 ppbv) the fraction of non-detects can be greater than 70%, and accounting for these values can decrease annual gridded averaged values by over 50% and make the distributions closer to what is expected based on surface station observations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".