A Federated Data Linkage Strategy to Support Population Health Research in Canada
Notice bibliographique
Résumé
ABSTRACT ObjectivesCanada has established a pan-Canadian cohort with over 300,000 volunteer participants aged 35-69, to support research on cancer and chronic disease. A key feature of the cohort is that participants have consented to link their cohort data with administrative datasets. This prospective cohort, representing nearly 1 in 50 Canadians in this age range will be followed for multiple decades, building a platform that supports access to timely, high-quality, data related to cancer and other chronic diseases, which will enable researchers to answer complex system questions and achieve better health outcomes for Canadians. ApproachA baseline “core” questionnaire was administered to participants capturing information on socio-demographics, economic characteristics, personal and familial history of diseases and lifestyle and health behaviours. A re-contact questionnaire is planned for 2016 to update baseline information and add depth to specific areas and capture changes over time. To realize the full potential for this cohort to support transformative research it is crucial to be able to link this data with other provincial datasets, such as cancer registries, hospital records and mortality data. For the most part, health data in Canada resides under the purview of health providers and, or government custodians in each of the provinces and territories. As such, an innovative federated data linkage strategy is required to link cohort data with health administrative data in each regions, adhering to existing privacy and regulatory requirements, while providing central access for researchers. ResultsTo date, 40% of all cohort participants have been linked with priority provincial administrative data. Each province has its own unique data linkage challenges, requiring unique customized solutions. By the end of 2017 we anticipate that the number of participants who will have had their data linked will increase to nearly 70%. The federated data linkage strategy and infrastructure offer an innovative approach that others can learn from; however to realize the full potential of the cohort and support transformative research partnerships and collaborations are required. ConclusionEfforts to organize resources and establish systems for data linkage and optimize data sharing and utilization in Canada are underway and include discussions with the provincial privacy commissioners, national and provincial and territorial data custodians and other thought leaders in the field.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,167 | 0,215 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,002 | 0,002 |
| Bibliométrie | 0,017 | 0,026 |
| Études des sciences et des technologies | 0,010 | 0,002 |
| Communication savante | 0,010 | 0,004 |
| Science ouverte | 0,008 | 0,018 |
| Intégrité de la recherche | 0,003 | 0,003 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,011 | 0,004 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».