A systematic review and scientific critique of methodology in modern urban heat island literature
Bibliographic record
Abstract
Abstract In the modern era of urban climatology, much emphasis has been placed on observing and documenting heat island magnitudes in cities around the world. Urban climate literature consequently boasts a remarkable accumulation of observational heat island studies. Through time, however, methodologists have raised concerns about the authenticity of these studies, especially regarding the measurement, definition and reporting of heat island magnitudes. This paper substantiates these concerns through a systematic review and scientific critique of heat island literature from the period 1950–2007. The review uses nine criteria of experimental design and communication to critically assess methodological quality in a sample of 190 heat island studies. Results of this assessment are discouraging: the mean quality score of the sample is just 50 percent, and nearly half of all urban heat island magnitudes reported in the sample are judged to be scientifically indefensible. Two areas of universal weakness in the literature sample are controlled measurement and openness of method : one‐half of the sample studies fail to sufficiently control the confounding effects of weather, relief or time on reported ‘urban’ heat island magnitudes, and three‐quarters fail to communicate basic metadata regarding instrumentation and field site characteristics. A large proportion of observational heat island literature is therefore compromised by poor scientific practice. This paper concludes with recommendations for improving method and communication in heat island studies through better scrutiny of findings and more rigorous reporting of primary research. Copyright © 2010 Royal Meteorological Society
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.370 | 0.599 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.007 | 0.004 |
| Bibliometrics | 0.021 | 0.027 |
| Science and technology studies | 0.002 | 0.010 |
| Scholarly communication | 0.009 | 0.008 |
| Open science | 0.005 | 0.005 |
| Research integrity | 0.007 | 0.003 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".