Bibliographic record
Abstract
Citizenship and Immigration Canada for providing their insights into the immigrant selection process and for their facilitation role in accessing the IMDB; Mike Abbott and Charles Beach at Queen’s Economics Department for their help with understanding the IMDB; and the Canadian Labour Market and Skills Researcher Network for its financial support and valuable feedback. We owe a special thanks to Tristan Cayn from CIC who worked tirelessly with us to run our code at Statistics Canada. There is growing international interest in a Canadian-style points system for selecting economic immigrants. Although existing points systems have been influenced by the human capital literature, the findings have traditionally been incorporated in an ad hoc way. This paper explores a formal method for designing a points system based on a human capital earnings regression for predicting immigrant economic success. The method is implemented for Canada using the IMDB, a remarkable longitudinal database that combines information on immigrants ’ characteristics at landing with their subsequent income performance as reported on tax returns. We demonstrate the feasibility of the method by developing an illustrative points system. We also explore how the selection system can be improved by incorporating additional information such as country-of-origin characteristics and intended occupations. We discuss what our findings imply for the debate about the relative merits of points- and employment-based systems for selecting economic immigrants. 1 1.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.044 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.018 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".