A Federated Data Linkage Strategy to Support Population Health Research in Canada
Bibliographic record
Abstract
ABSTRACT
 ObjectivesCanada has established a pan-Canadian cohort with over 300,000 volunteer participants aged 35-69, to support research on cancer and chronic disease. A key feature of the cohort is that participants have consented to link their cohort data with administrative datasets. This prospective cohort, representing nearly 1 in 50 Canadians in this age range will be followed for multiple decades, building a platform that supports access to timely, high-quality, data related to cancer and other chronic diseases, which will enable researchers to answer complex system questions and achieve better health outcomes for Canadians.
 ApproachA baseline “core” questionnaire was administered to participants capturing information on socio-demographics, economic characteristics, personal and familial history of diseases and lifestyle and health behaviours. A re-contact questionnaire is planned for 2016 to update baseline information and add depth to specific areas and capture changes over time. To realize the full potential for this cohort to support transformative research it is crucial to be able to link this data with other provincial datasets, such as cancer registries, hospital records and mortality data.
 For the most part, health data in Canada resides under the purview of health providers and, or government custodians in each of the provinces and territories. As such, an innovative federated data linkage strategy is required to link cohort data with health administrative data in each regions, adhering to existing privacy and regulatory requirements, while providing central access for researchers.
 ResultsTo date, 40% of all cohort participants have been linked with priority provincial administrative data. Each province has its own unique data linkage challenges, requiring unique customized solutions. By the end of 2017 we anticipate that the number of participants who will have had their data linked will increase to nearly 70%. The federated data linkage strategy and infrastructure offer an innovative approach that others can learn from; however to realize the full potential of the cohort and support transformative research partnerships and collaborations are required.
 ConclusionEfforts to organize resources and establish systems for data linkage and optimize data sharing and utilization in Canada are underway and include discussions with the provincial privacy commissioners, national and provincial and territorial data custodians and other thought leaders in the field.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.041 | 0.015 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.007 | 0.010 |
| Open science | 0.017 | 0.005 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".