Building and occupant characteristics as predictors of temperature-related health hazards in American homes
Bibliographic record
Abstract
• Extreme temp. exposure often happens at home, but role of buildings is unclear • We built machine learning models to predict temperature-related illness in US homes • Compared model performance with different variables from US household survey data • Including building-related variables improves model accuracy, recall, and precision • Results aid public health planning to mitigate temperature-related health hazards Many cities and regions are making significant investments towards planning for extreme temperature and in particular extreme heat. A heat vulnerability index (HVI) is a metric to track spatial variation in extreme temperature risk to target mitigation interventions. Most HVIs focus on demographic characteristics, which generally relate to vulnerability, and lack information about the building stock, which mediate the occupant's exposure to extreme temperatures. In this study, we use the Energy Information Administration's (EIA) Residential Energy Consumption Survey (RECS) to estimate prevalence of temperature-related illness in the United States and develop machine learning models using climate, demographic, and building characteristics to predict them. Temperature-related illness affects approximately 2 million households annually, around 1% of the total population. The models we developed predict temperature-related illness with up to 85% accuracy. The most important feature is energy insecurity, which describes the household's ability to maintain and operate heating, ventilation, and air conditioning (HVAC) systems. Our results offer guidance for municipalities to improve data collection, enabling them to better identify at-risk households and strategize resources for short-term and long-term interventions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".