Bibliographic record
Abstract
To generate a climate data set, temperature data collected at the Earth's surface must be adjusted to remove non-climatic effects such as urbanization and measurement discontinuities. Some studies have shown that the post-1980 spatial pattern of temperature trends over land in prominent climate data sets is strongly correlated with the spatial pattern of socioeconomic development, implying that the adjustments are inadequate, leaving a residual warm bias. This evidence has been disputed on three grounds: spatial autocorrelation of the temperature field undermines significance of test results; counterfactual experiments using model-generated data suggest such correlations have an innocuous interpretation; and different satellite covariates yield unstable results. Somewhat surprisingly, these claims have not been put into a coherent framework for the purpose of statistical testing. We combine economic and climatological data sets from various teams with trend estimates from global climate models and we use spatial regressions to test the competing hypotheses. Overall we find that the evidence for contamination of climatic data is robust across numerous data sets, it is not undermined by controlling for spatial autocorrelation, and the patterns are not explained by climate models. Consequently we conclude that important data products used for the analysis of climate change over global land surfaces may be contaminated with socioeconomic patterns related to urbanization and other socioeconomic processes. Research Supported by Social Sciences and Humanities Research Council of Canada Grant Number 430002.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.034 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.018 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.011 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".