Racial Segregation and Discrimination: Evidence from the Rental Housing Market
Bibliographic record
Abstract
Rental housing discrimination can take many forms. The most visible potential result of discrimination is the segregation of neighborhoods by race and ethnicity. The more subtle form of discrimination examined in this study is to charge different prices to different groups. This form of discrimination is possible if, among other things, owners of housing units believe unobserved factors important to the return on their investment are correlated with race or ethnicity, or if landlords have a taste for discrimination as described by Becker (1957). Regardless of the reason why landlords desire to charge different prices to different groups, theory suggests that these differences will be more pronounced in tighter housing markets, where the cost of discrimination to the landlord is lower. The study proposed here is an extension of the use of hedonic methods to detect discrimination using a much richer set of data than what has been available to previous authors (Follain and Malpezzi (1981), Kiel and Zabel (1995), and Myers (2004) for example) to answer three policy relevant questions related to the variations in access to housing across race and ethnicity: 1) Do minorities pay more for equal quality housing to live in majority neighborhoods? 2) Do minorities pay more for equal quality housing to live in areas with less concentrations of poverty? and 3) Does the tightness of the housing market effect the ability of landlords to charge different rents for equal quality housing based on race and ethnicity? The last question appears to have been largely ignored in previous empirical studies of housing discrimination. We make a considerable effort to include a comprehensive set of covariates in the hedonic equation that, due to data constraints, were not considered in previous studies. First, our primary source of data includes over 450,000 observations on rental housing units with more than 70 detailed questions about the units type, condition, and neighborhood attributes. Second, these data can be linked with data from the 2000 Decennial Census providing additional detailed characteristics about the neighborhood. Finally, information about credit history (at the census tract level) has been provided by Equifax, one of the three national credit reporting agencies. Although it is aggregated at the census tract level, the credit data is notably rich.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.012 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.009 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".