Estimating biases and error variances through the comparison of coincident satellite measurements
Bibliographic record
Abstract
A framework for the statistical comparison of six coincident remote sounding measurements is presented, which distinguishes between additive and multiplicative biases. The relationship between multiplicative bias and error variance is explored, and three methods are proposed for producing sets of values for three comparison variables: the multiplicative bias, and the error variance for each of two instruments. We illustrate and compare the three methods through the comparison of coincident measurements of the relatively long‐lived stratospheric species O3, N2O, and HNO3 from two independent measurement sets: version 2.2 retrievals (with updated O3) from the Atmospheric Chemistry Experiment‐Fourier transform spectrometer onboard SCISAT‐1, and version 1.51 retrievals from the Earth Observing System Microwave Limb Sounder onboard Aura. We find that multiplicative bias between the two measurement sets, compared on a common vertical grid, is significant at some heights for O3 and N2O, and for all heights tested for HNO3. The most realistic estimates of measurement error are produced by a method which incorporates a third correlative data set into the analysis. Using this method, estimated error standard deviations (SDs) are comparable between the two instruments for O3 measurements, and are less than 10% of the mean measurement value between approximately 100 and 1 hPa. ACE N2O measurements are consistent with a 10% error SD at all heights tested, although the uncertainty of the estimates is large at heights above 5 hPa. Estimated MLS N2O error SDs are comparable with those for ACE in the lower stratosphere, but increase steeply with height. For HNO3, estimated error SDs are approximately 10% between 70 and 10 hPa for both instruments. At heights above 10 hPa and below 100 hPa, estimated ACE errors are significantly smaller than those for MLS.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.019 | 0.082 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".