Which human capital characteristics best predict the earnings of economic immigrants
Bibliographic record
Abstract
While there is an extensive literature on immigrant entry earnings in Canada, there is a lack of knowledge, when predicting immigrant earnings, on the relative importance of various human capital factors, such as language, work experience, age, and education. This paper addresses two questions. First, what is the relative importance of such observable human capital factors when predicting the earnings of economic immigrants (principal applicants) who are selected through the points system? Second, does the relative importance of these factors vary in the short, intermediate, and long term? Using the Longitudinal Immigration Database, the study finds that the predictive power of immigrant characteristics (measured at landing) changes with years spent in Canada. Official-language background at landing, and Canadian work experience before immigration, are the best predictors of annual earnings in the first two years after landing for economic immigrants (principal applicants). However, educational attainment at landing and age at landing (a proxy for foreign work experience) are the best predictors of longer-term earnings (10 to 11 years after landing). Some interaction effects are also important. The predictive power of education and age (in part a proxy for foreign work experience) is influenced by their interaction with official-language skills and Canadian work experience. The earnings advantage of higher education is much larger among principal applicants who have strong rather than weak official-language skills. Immigrants whose mother tongue is English or French do not experience a significant negative effect of age on earnings. Finally, many factors beyond those studied here affect immigrant earnings. The predictive power of regression models could be increased with improved data sources.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".