Metrics of Urbanicity and Rurality in US-Based Epidemiologic Studies of Ambient Temperature and Health: A Scoping Review
Bibliographic record
Abstract
BACKGROUND: The impacts of environmental health risk factors, including temperature, vary across urban and rural areas. Application of different metrics of rurality and urbanicity can yield different risk characterizations. We aimed to identify, describe, and quantify how urban/rural metrics are used in epidemiologic studies of ambient temperature and health across the United States (US). METHODS: Using PubMed and Scopus, we identified epidemiologic studies published between January 2010 and March 2025 that examined ambient temperature and health in the US and included a defined, quantitative metric of urbanicity/rurality. Titles, abstracts, and full texts were evaluated by two independent reviewers. Data from included studies were extracted using a predetermined tool. RESULTS: Of the 11,013 studies resulting from our search, 36 were included. We identified 23 metrics drawing from 10 data sources. The most frequently used metrics were population density and size from the US Census (n = 11 studies). Other metrics reflected connectivity and proximity to surrounding areas, such as the US Census’s Urban-Rural Classification (n = 7 studies), and the US Department of Agriculture’s Rural-Urban Commuting Area Codes (n = 4 studies) and Rural-Urban Continuum Codes (n = 2 studies). Additional metrics captured features related to the natural environment, built environment, and employment. Many studies did not provide a rationale for metric selection. DISCUSSION: Urbanicity and rurality metrics have moved beyond population size and density to include other features. Providing rationales for choice of metric or the differential vulnerability or adaptive capacity captured by the metric could bolster understanding of urban-rural differences in the impact of temperature on health.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.053 | 0.269 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.006 | 0.009 |
| Bibliometrics | 0.048 | 0.048 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.007 | 0.008 |
| Open science | 0.003 | 0.005 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".