Bibliographic record
Abstract
There is growing belief in many developed countries, including Canada, that the large influx of the foreign-born population increases crime. Despite the heated public discussion, the immigrant-crime relationship is understudied in the literature. This paper identifies the causal linkages between immigration and crime using panel data constructed from the Uniform Crime Reporting Survey and the master files of the Census of Canada. This paper distinguishes immigrants by their years in Canada and defines three groups: new immigrants, recent immigrants and established immigrants. An instrumental variable strategy based on the historical ethnic distribution is used to correct for the endogenous location choice of immigrants. Two robust patterns emerge. First, new immigrants do not have a significant impact on the property crime rate, but with time spent in Canada, a 10% increase in the recent-immigrant share or established-immigrant share decreases the property crime rate by 2% to 3%. Neither underreporting to police nor the dilution of the criminal pool by the addition of law-abiding immigrants can fully explain the size of the estimates. This suggests that immigration has a spillover effect, such as changing neighbourhood characteristics, which reduces crime rates in the long run. Second, IV estimates are consistently more negative than their OLS counterparts. By not correctly identifying the causal channel, OLS estimation leads to the incorrect conclusion that immigration is associated with higher crime rates.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.009 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.011 |
| Science and technology studies | 0.004 | 0.002 |
| Scholarly communication | 0.002 | 0.000 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".