How integration of the federal Indian Register has enhanced First Nations-specific analysis of ICES data
Bibliographic record
Abstract
IntroductionIn Ontario, First Nations are increasingly seeking population-level data about the health of their citizens. However, First Nations people are not readily identified in standard health administrative data and indirect strategies, such as the use of on-reserve addresses, are limited in scope and validity. Objectives and ApproachThe Chiefs of Ontario entered into a Data Governance Agreement with the Institute for Clinical Evaluative Sciences (ICES) that enabled the linkage of the federal Indian Register (IR) to data at ICES. This study examined the impact of the IR linkage on First Nations population estimates and location of residence, measured by postal code or residence code. Overall, and for each First Nation community in Ontario, we compared First Nations population estimates from the ICES data with and without the IR linkage to estimates available from Indigenous and Northern Affairs Canada (INAC). ResultsWithout the IR, using only Ontario residence codes or postal codes that were unique to a given community, 62,242 individuals were identified as living in First Nations communities. This is approximately 30% lower than the current INAC on-reserve population estimate of 92,234 for First Nations communities in Ontario. Adding the IR allowed the use of non-unique postal codes as well, resulting in the identification of an additional 15,183 First Nations individuals. It also allowed the identification of over 113,000 First Nations individuals who live outside of First Nations communities, especially in urban areas. Finally, the combination of residence information and the IR permits communities to identify their registered member living within and outside their communities. Conclusion/ImplicationsUsing the IR in combination with geographic residence information, made possible through the Data Governance Agreement signed between Chiefs of Ontario and ICES, will provide First Nations communities with more accurate and complete population estimates, which is key to the production of useful and relevant First Nations-specific health research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.180 | 0.464 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.009 | 0.020 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.010 | 0.006 |
| Open science | 0.003 | 0.008 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.007 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".