What Do We Know About Location Affordability in U.S. Shrinking Cities?
Bibliographic record
Abstract
In late 2013, the Department of Housing and Urban Development (HUD) launched the Location Affordability Index (LAI) portal. Their dataset uses models to estimate typical amount households spend on housing and transportation at the block group level, and calculates “H + T Affordability,” the percent of household income spent on these items. In our previous research, we analyzed 81 shrinking cities to determine how location affordability differs across various neighborhoods. Our results suggest that households in declining neighborhoods, as compared to stable or redeveloping neighborhoods, face the greatest H + T affordability challenges in shrinking cities. Furthermore, in declining neighborhoods, virtually all of the additional affordability challenges encountered can be accounted for by differences in transportation affordability rather than housing. Since there is virtually no research to either validate or suggest bias in the LAI data, and a declining neighborhood in a shrinking city presents both a relatively common yet entirely dissimilar context to the norm, we feel that this data should be carefully calibrated to, and tested for, this setting to ensure that appropriate and efficient policy follow. In this report, we present the results of two research phases: a strictly quantitative first stage in which the LAI is disassembled and reassembled, taking stock of assumptions, methods, and accuracy. This work finds that the LAI generally over-estimates housing costs, but more for renters and more in metropolitan areas. Estimations of transportation cost burdens are built largely from unreliable data and using models which cannot be replicated, leading us to conclude that these cost burden estimates may not be reliable. In the second research phase, which is survey-based, we gather household-level data from 12 Census tracts in Cleveland, Ohio, to estimate household housing and transportation costs and cost burdens and gain a clearer sense of budget trade-offs where costs are unaffordable. The survey results support the first research stage. The LAI over-estimates housing costs for these neighborhoods by approximately 17% and transportation costs by approximately 126%, meaning the LAI estimates for transportation are not reliable. By and large, households trade essentials, investing, and paying bills (at all or in full) to cover their costs in housing and transportation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.027 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.013 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.003 | 0.006 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".