Supplementary material to "How well do Earth System Models reproduce observed aerosol changes during the Spring 2020 COVID-19 lockdowns?"
Bibliographic record
Abstract
S1 Observational Uncertainties and Sampling DifferencesThis work is based on comparisons between observed and simulated anomalies of AOD.In order to draw meaningful conclusions from these comparisons we need to assess how much of the difference between our datasets comes from differences in sampling, and how much can be attributed to uncertainties or biases in the products themselves.Here we assess the role of sampling differences in explaining the considerable spread between AOD estimates from different satellites, which can differ by 20-30% (Table S1).Our main analysis used monthly and regional mean data products.For the simulations, these means are spatiotemporally complete; the observed means use retrievals obtained at the satellite's particular overpass time, in clear-sky conditions (for the passive sensors), when the retrieval was successful and not prevented by a myriad of potential limitations such as sun glint or complex terrain.Here we conduct a systematic intercomparison between pairs of observational data products, with each pairing selected to isolate the effect of a particular sampling effect.Discrepancies that cannot be attributed to sampling differences are then used to estimate the degree of uncertainty in our observed AOD responses.Throughout, figures that do not require gridcellto-gridcell comparisons (spatial-mean timeseries, histograms) are generated at the products' native resolutions, and figures that do require gridcell-to-gridcell comparisons (scatter plots, calculation of correlation between datasets) are generated on fields that have been interpolated to a common 1 • ×1 • resolution. S1.1 Temporal SamplingWe first assess the effects of temporal sampling on our comparison between satellite observations, which are sampled only at a particular overpass time, and model simulations, which are integrated over the full diurnal cycle.Figure S1 compares threehourly, daily, and monthly AOD values for MAM 2004 from a sample CanAM5 simulation.Timeseries plot the evolution of AOD over the three-month period in each of our analysis regions, with blue, orange, and green lines corresponding to the different averaging timescales.A dotted black line indicates the diurnally-integrated, MAM-mean AOD.For all regions except the Northern Hemisphere, blue horizontal lines additionally show the mean of AOD for MAM 2004 sampled only at our 1 Table S1.Mean AOD over reference period (2015)(2016)(2017)(2018)(2019) from different satellites, and the spread in these estimates.The relative spread is calculated with respect to the mean of the highest and lowest AOD estimates in each region. RegionMODIS Aqua MISR CALIOP ASN ACROS-C absolute spread relative spread N. Hemis 0.231 0.190 0.178 0.181 0.053 26% E.China 0.483 0.331 0.445 0.451 0.152 30% India 0.402 0.365 0.459 0.477 0.112 27% Europe 0.157 0.125 0.127 0.124 0.033 23% Table S2.Mean AOD values obtained from 3-hourly CanAM output, when averaged over all of MAM 2004 ("full day") and when sampled only at satellite overpass times.Values in parentheses indicate the percent change in AOD when sampled at a particular overpass time as compared to the diurnally-integrated value.These relative differences are an order of magnitude smaller than the spread bewteen region-mean AOD from different satellites (Table S1).Region Full Day Aqua Day Overpass Aqua Night Overpass Terra Day Overpass N. Hemis 0.311 ---E.China 0.503 0.488 (-3.0%) 0.513 (+2.0%) 0.499 (-0.8%)India 0.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.002 | 0.011 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.008 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".