Cancer and the healthy immigrant effect: a statistical analysis of cancer diagnosis using a linked Census-cancer registry administrative database
Bibliographic record
Abstract
BACKGROUND: A large volume of research has been published on both the socio economic and demographic determinants of cancer and on the health of immigrants and minority groups. Yet because of data limitations, little research examines differences in the occurrence of cancer incidence between immigrants and non-immigrants and among immigrants defined by region of birth and time in the host country. In particular it is not known whether a healthy immigrant effect is present for cancer and if so, whether this advantage is lost with additional years of residence in the host country. METHODS: This paper uses a large data file from Statistics Canada that links Census information on immigrant status, socioeconomic status including educational attainment, and other person-level information with administrative data on cancer and mortality over a continuous 13 year period of observation. It estimates discrete and continuous time duration models to identify differences in cancer diagnosis by immigrant subgroup after controlling for a variety of potential confounders. Differences in historical smoking behavior are not observable at the individual level in the dataset but are accounted for indirectly using various methods. RESULTS: Results in general confirm the existence of a healthy immigrant effect for cancer in that, overall, recent immigrants to Canada are significantly less likely than otherwise comparable non-immigrant Canadians to be diagnosed with any cancer and the most common forms of cancer by site. As well, this gap appears to decline with additional years in Canada for immigrant men and women, eventually converging to Canadian-born levels. Differentiating among immigrant subgroups by period of arrival and country of birth reveals significant variation across immigrant subgroups, with immigrant men and women from developing countries typically having a lower likelihood of being diagnosed with cancer than immigrants from the US, UK and continental Europe. As well, controlling for immigrant heterogeneity this way weakens the conclusion that the gap narrows with years in Canada. Immigrant men overall continue to exhibit convergence to Canadian-born levels for diagnosis of any cancer and for prostate cancer, while immigrant women exhibit narrowing over time only for breast cancer. Although smoking behavior is not directly observed, controlling for subgroup-specific lifetime smoking behavior using survey data has only a relatively minor effect on the estimated differences. CONCLUSIONS: The specificity of the results by cancer type, gender, immigrant status and ethnicity provides useful guidance for future research by helping to narrow the possible channels through which social and economic characteristics may be affecting cancer incidence.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.027 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.003 | 0.006 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".