Building a Pan-Canadian Real World Health Data Network
Bibliographic record
Abstract
BackgroundIn December 2017 the Canadian Institutes of Health Research (CIHR) issued a request for proposals to develop a pan-Canadian health data platform. This platform will enable cross-jurisdictional research by facilitating the use of rich provincial and national data and ensure engagement with patients and specific populations including Indigenous partners. Academics and policy makers from across Canada operating under the banner of the Pan-Canadian Real-World Health Data Network (PRHDN) have joined forces to address this call.
 ObjectivesCreate national infrastructure that is built once then made available for research, benchmarking, performance monitoring, multi-jurisdictional evaluations and inter-jurisdictional comparisons to address pressing health and social policy problems in Canada.
 MethodsOur approach will address several issues including creating significant efficiencies in data access, streamlining cross provincial/ territorial ethics and access approvals, establishing standards for data and methods harmonization and providing innovative and privacy-conscious solutions to data access and use. The presentation will focus on the plan to create harmonized common data, algorithms and analytic protocols, and link administrative data to electronic medical records and clinical trials to create an integrated and documented infrastructure for pan-Canadian studies. Comparisons to PopMedNet and the Sentinel Initiative in the US will be made.
 ConclusionProvincial centres across Canada hold rich sources of health and social data that are linkable at the person-level. With the exception of standardized data managed by the Canadian Institute for Health Information (CIHI), these data are often not comparable from one province to another, thereby limiting use to single-province studies. There is growing interest in Canada in creating an environment that would enable cross-jurisdictional data sharing and analysis’ and in sharing experiences to make effective use of linkable administrative data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.004 | 0.000 |
| Scholarly communication | 0.000 | 0.003 |
| Open science | 0.005 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".