Bibliographic record
Abstract
The earth’s near surface air temperature is important in a variety of applications including for quantifying global warming. We analyze 6 monthly series of atmospheric temperatures from 1880 to 2012, each produced with different methodologies. We first estimate the relative error by systematically determining how close the different series are to each other, the error at a given time scale is quantified by the root mean square fluctuations in the pairwise differences between the series as well as between the individual series and the average of all the available series. By examining the differences systematically from months to over a century, we find that the standard short range correlation assumption is untenable, that the differences in the series have long range statistical dependencies and that the error is roughly constant between 1 month and one century—over most of the scale range, varying between ±0.03 and ±0.05 K. The second part estimates the absolute measurement errors. First we make a stochastic model of both the true earth temperature and then of the measurement errors. The former involves a scaling (fractional Gaussian noise) natural variability term as well as a linear (anthropogenic) trend. The measurement error model involves three terms: a classical short range error, a term due to missing data and a scale reduction term due to insufficient space–time averaging. We find that at 1 month, the classical error is ≈±0.01 K, it decreases rapidly at longer times and it is dominated by the others. Up to 10–20 years, the missing data error gives the dominate contribution to the error: 15 ± 10% of the temperature variance; at scales >10 years, the scale reduction factor dominates, it increases the amplitude of the temperature anomalies by 11 ± 8% (these uncertainties quantify the series to series variations). Finally, both the model itself as well as the statistical sampling and analysis techniques are verified on stochastic simulations that show that the model well reproduces the individual series fluctuation statistics as well as the series to series fluctuation statistics. The stochastic model allows us to conclude that with 90% certainty, the absolute monthly and globally averaged temperature will lie in the range −0.109 to 0.127 °C of the measured temperature. Similarly, with 90% certainty, for a given series, the temperature change since 1880 is correctly estimated to within ±0.108 of its value.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".