Cultural and contextual relevance of the Indigenous data in the Canadian longitudinal study on aging
Bibliographic record
Abstract
OBJECTIVES: The CLSA is a national data platform for aging research that used epidemiology-based sampling methods and explicitly excluded people living on First Nations Reserves and other provincial First Nations settlements as possible CLSA participants. As such, the CLSA research approach did not use Indigenous community engagement. Nevertheless, the CLSA sample includes a sizeable subsample of participants who self-identified as First Nations, Métis, or Inuit. This project seeks to describe the self-identified Indigenous subsample of the CLSA from the baseline data collection and interpret that description with the aid of an Elder Advisory Circle. METHODS: We conducted a descriptive analysis of the self-identified Indigenous subsample of the CLSA from the baseline data collection. The analysis was presented to an Elder Advisory Circle for consultation. RESULTS: The lack of community-engaged approaches to Indigenous research and sampling approaches appears to have resulted in a sociodemographic profile of older Indigenous Peoples that does not match the lived experience of the Elder Advisory Circle and contrasts with other data available on Indigenous Peoples in Canada. We feel the existing CLSA data does not reflect the sociodemographic profile of older Indigenous Peoples. CONCLUSION: We use this community consultation process to provide recommendations for the appropriate use of the Indigenous-identified data in the CLSA, and we conclude by recommending great caution when using the data from the Indigenous subsample in the CLSA data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.102 | 0.189 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.006 | 0.011 |
| Science and technology studies | 0.027 | 0.012 |
| Scholarly communication | 0.009 | 0.003 |
| Open science | 0.004 | 0.009 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".