Bibliographic record
Abstract
This dissertation examines immigrants (to Canada) assimilation problems from a perspective of imperfect human capital transferability. Chapter 2 discusses how much of the immigrant wage gap can be explained by the undervaluation of foreign human capital (education and work experience). The identification of the human capital source (using information available in the 2006 Canadian census) can explain up to 70% of the native-immigrant wage gap. The foreign-born dummy coefficient goes from around -11% to close to -3%. Education acquired in Asia tends to be valued less than education from South America, Africa and East Europe; which in turn is less valued than education from Oceania, the U.S. and the rest of continental Europe. Studying in the UK consistently appears more beneficial than studying in Canada. When incorporating country of origin fixed effects, the different specifications visibly reduce the heterogeneity of country coefficients. The reduction is sizeable for Pakistan, India, China and the Philippines; though their coefficients remain negative. A smaller reduction for Europe, South-East Asia, Hong Kong and the US drives their coefficients close to zero. The UK country of origin dummy has the only persistently positive coefficient. Chapter 3 describes the occupational assimilation process of 2000-2001 immigrants in their first four years. The results show that those with high levels of education experience a more significant decline in their first occupation. Education, though, has a positive and significant effect on occupational improvement; which reduces the size and significance of the negative effect of education on the second occupational gap. It, however, does not change its sign. The same pattern is observed when analyzing occupational gaps through time. Chapter 4 focuses on immigrants' English proficiency improvement. Overall, immigrants show relatively small improvements in language proficiency in the first four years in Canada. Still, those arriving under the family immigrant category with an intermediate or advanced level are less likely to improve and more likely to decrease their English proficiency. Human capital variables (age and education) are also consistently relevant for English proficiency improvement.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.008 | 0.005 |
| Scholarly communication | 0.004 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.014 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".