The Nascent Pan-Canadian Real-world Health Data Network (PRHDN)
Bibliographic record
Abstract
ABSTRACT
 ObjectiveCanada’s large single payer health systems have created provincial centres with rich, population-wide health and social data holdings that are linkable at the person level. Complementary federal data include a subset of standardized and enriched administrative healthcare data at the Canadian Institute for Health Information, Statistics Canada’s extensive array of survey- and census-based social and geography-based information, and large national datasets that are being created through pan-Canadian initiatives. Unfortunately, these data assets are rarely combined in multi-province or pan-Canadian studies, often because data are not directly comparable from one province to another, or cannot be shared due to legislative or other barriers. There is now growing interest in enabling multi-province studies — even when data are not directly comparable — and in sharing experiences to make more effective use of linked or linkable administrative data across Canada.
 ApproachNine provincial and national organizations have created a detailed Implementation Plan for the new PRHDN distributed data network. Without requiring that record-level data leave provincial boundaries, the PRHDN will create shared core research data infrastructure: (i) validated algorithms that implement case definitions applicable across provinces, (ii) harmonized common data and (iii) common analytic protocols. The PRHDN will also establish complementary infrastructure including dedicated personnel to assist researchers and decision makers, joined-up training and capacity building sessions, opportunities to share learning and experience related to linking new datasets such as electronic medical records and omic datasets, and forums for knowledge translation and exchange with decision makers.
 ResultsThe PRHDN is already bringing together expertise from across Canada as researchers, decision makers and data custodians begin to identify opportunities for enhanced use of health data in Canadian research, policy making and practice. A simple PRHDN website has been created (https://www.prhdn.ca/) and more than 200 researchers and policy/decision makers have become members of the PRHDN consortium. Surveys of consortium members are identifying priorities for the first algorithm validation work (to-date the top four priorities are mental health, cardiovascular disease, diabetes and respiratory disease) and the PRHDN Leads Team is functioning as a decision-making body, e.g., in discussions with Statistics Canada.
 ConclusionsWorking together, provincial and national organizations across Canada have identified concrete steps that can be taken to enable multi-province and pan-Canadian studies based on administrative data within one year of the start of funding. The PRHDN Leads Team is currently discussing the PRHDN vision and Implementation Plan with potential funders.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.019 | 0.000 |
| Scholarly communication | 0.001 | 0.004 |
| Open science | 0.011 | 0.002 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".