Supplementary material to "How well do Earth System Models reproduce observed aerosol changes during the Spring 2020 COVID-19 lockdowns?"
Bibliographic record
Abstract
Observational Uncertainties and Sampling DifferencesThis work is based on comparisons between observed and simulated anomalies of AOD.In order to draw meaningful conclusions from these comparisons we need to assess how much of the difference between our datasets comes from differences in sampling, and how much can be attributed to uncertainties or biases in the products themselves.Here we assess the role of sampling differences in explaining the considerable spread between AOD estimates from different satellites, which can differ 5 by 20-30% (Table S1).Our main analysis used monthly and regional mean data products.For the simulations, these means are spatiotemporally complete; the observed means use retrievals obtained at the satellite's particular overpass time, in clear-sky conditions (for the passive sensors), when the retrieval was successful and not prevented by a myriad of potential limitations such as sun glint or complex terrain.Here we conduct a systematic intercomparison between pairs of observational data products, with each pairing 10 selected to isolate the effect of a particular sampling effect.Discrepancies that cannot be attributed to sampling differences are then used to estimate the degree of uncertainty in our observed AOD responses.Throughout, figures that do not require gridcellto-gridcell comparisons (spatial-mean timeseries, histograms) are generated at the products' native resolutions, and figures that do require gridcell-to-gridcell comparisons (scatter plots, calculation of correlation between datasets) are generated on fields that have been interpolated to a common 1 • ×1 • resolution. 15 S1.1 Temporal SamplingWe first assess the effects of temporal sampling on our comparison between satellite observations, which are sampled only at a particular overpass time, and model simulations, which are integrated over the full diurnal cycle.Figure S1 compares threehourly, daily, and monthly AOD values for MAM 2004 from a sample CanAM5 simulation.Timeseries plot the evolution of AOD over the three-month period in each of our analysis regions, with blue, orange, and green lines corresponding to the 20 different averaging timescales.A dotted black line indicates the diurnally-integrated, MAM-mean AOD.For all regions except the Northern Hemisphere, blue horizontal lines additionally show the mean of AOD for MAM 2004 sampled only at our
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.016 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.002 | 0.004 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.810 | 0.256 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".