Investigating the Spatial Patterns of Income Inequality and Crime: Applications of Spatiotemporal Data Analysis Techniques
Bibliographic record
Abstract
Income inequality and crime are two social problems concerning many nations around the world. Such social issues can have detrimental effects on society and are interconnected, meaning that high crime rates may be associated with high levels of income inequality. Spatial data analysis of income inequality and crime can potentially aid in planning inequality reduction and crime prevention measures. A variety of studies have been conducted to explore the spatial patterns of income inequality and crime, but there are still research gaps and uncertainties that exist. First, while income inequality has been analyzed at various spatial scales, there is a lack of research conducted at the small area level within cities and neighbourhoods. Second, while criminology theories such as rational choice theory indicate a positive association between spatial patterns of income inequality and crime, empirical studies have produced inconsistent and sometimes contradictory results. Third, while some studies suggest that the spatial and temporal dimensions of crime are inseparable, research on the spatiotemporal dimensions of crime is limited compared to purely spatial studies. This thesis aims to investigate the spatial variability of income inequality, the relationship between income inequality and crime, and the spatiotemporal variation of crime between business days and non-business days at the small area level. This thesis adopts a manuscript-style format consisting of three papers. \nThe first paper adopts an exploratory spatial data analysis approach to examine the spatial patterns of income inequality in the City of Toronto at two spatial scales: census tract and dissemination area. Noteworthy locations of within-area income inequality, represented by the Gini coefficient in each area, and across-area income inequality, represented by the median income disparities between different areas, are identified at each spatial scale. This paper also recognizes discrepancies in spatial patterns between the two spatial scales of analysis, since dissemination areas tend to capture more detailed local variation. The issue of scale can be attributed to the modifiable areal unit problem, where different spatial data aggregation units may lead to different statistical results and conclusions. \nThe second paper applies non-spatial and spatial regression models using frequentist and Bayesian modelling frameworks to explore the impacts of within-area and across-area income inequality on five major crime types in the City of Toronto at the census tract and dissemination area scales. The use of spatial regression models improves the model fit in both frequentist and Bayesian frameworks. The Bayesian shared component model accounts for the interactions between crime types and further enhances model performance. Results obtained from the best-fitting frequentist and Bayesian models are inconsistent but do not conflict in terms of the relationship between crime and income inequality, where within-area income inequality generally increases major crime rates and across-area income inequality has varying effects depending on the crime type and spatial unit of analysis. \nThe third paper investigates the small-area spatiotemporal variation of five major crime types between business days and non-business days using Bayesian modelling. The study area is Old Toronto, a district of high political and economic activity within the City of Toronto. The results of this paper indicate that non-business days tend to have higher risks of assault and robbery compared to business days, but the overall changes in auto theft, break and enter, and theft over $5,000 are insignificant. For each crime type, the local temporal trends in small areas vary across the study region. Although locations of significant temporal trends are identified for every crime type, crime hot spots generally do not differ between business days and non-business days. Nevertheless, some areas that are considered to be hot spots of assault, robbery or auto theft in both time periods have significantly higher crime risks on non-business days compared to business days. Additionally, sociodemographic characteristics (e.g., low income, residential instability) and built environment factors (parks and business areas) are found to be significantly associated with the spatial patterns of crime, while built environments (schools, parks, and business areas) also explain some local temporal variations of crime.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.033 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.010 | 0.019 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.004 | 0.002 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".