On sampling bias adjustment for sparsely observing satellite instruments for the example of carbonyl sulfide (OCS)
Bibliographic record
Abstract
Abstract. When computing climatological averages of atmospheric trace gas mixing ratios obtained from satellite-based measurements, sampling biases arise if data coverage is not uniform in space and time. Complete homogeneous spatio-temporal coverage is essentially impossible to achieve. Solar occultation measurements, by virtue of satellite orbits and the requirement of direct observation of the sun through the atmosphere, result in particularly sparse spatial coverage. In this study, a method is presented to adjust for such sampling biases when calculating climatological means. The method is demonstrated using carbonyl sulfide (OCS) measurements at 16 km altitude from the ACE-FTS (Atmospheric Chemistry Experiment Fourier Transform 15 Spectrometer). At this altitude, OCS mixing ratios show a steep gradient between the poles and equator. ACE-FTS measurements, which are provided as vertically resolved profiles, and integrated stratospheric OCS columns are used in this study. The bias adjustment procedure requires no additional observations other than the satellite data product itself and is expected to be generally applicable when constructing climatologies of long-lived tracers from sparsely and heterogeneously sampled satellite data. In a first step of the adjustment procedure, a regression model is used to fit a 2-D surface to all available ACE-FTS OCS measurements as a function of day-of-year and latitude. The regression model fit is used to calculate an adjustment factor, 20 which is then used to adjust each measurement individually. The mean of the adjusted measurement points of a chosen spatio-temporal frame is then used as the bias-free climatological value. When applying the adjustment factor to seasonal averages in 30° zones, the maximum spatio-temporal sampling bias adjustment was 11 % for OCS mixing ratios at 16 km and 5 % for the stratospheric OCS column. The adjustments were validated against the much denser and more homogeneous OCS data product from the limb-sounding MIPAS (Michelson Interferometer for Passive Atmospheric Sounding) instrument, and both the direction and sign of the adjustments were in agreement with the adjustment of the ACE-FTS data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.015 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".