One for all, all for one - Establishing a corporate linkage methodology for integrated health analytics in Canada
Bibliographic record
Abstract
ABSTRACTObjectivesTo describe the successes, challenges and trade-offs encountered when establishing a standard patient linkage methodology for linking pan-Canadian health data across the continuum of care.
 ApproachHealth facilities, regional health authorities and ministries of health across Canada need access to timely, integrated health information across the continuum of care in order to make effective decisions to manage and improve their health systems. To better meet these needs, our organization has embarked on an initiative to integrate our web-based analytic tools. A requirement for this integration is the establishment of a corporate linkage standard.
 The scope is to enable linkage across a dozen patient-level data holdings including data from acute care, long term care, home care, rehabilitation, mental health, pharmaceuticals and registries for joint replacements and organ transplants. In Canada, provinces and territories issue jurisdiction-specific health care numbers (HCN) to their residents.
 A working group was established to review existing methodologies and to define a standard linkage. To gain support from the various groups in our organization to adopt the proposed standard, the right balance was sought between accuracy, timeliness, ease of implementation as well as availability and quality of data elements across our data holdings. In addition, an accurate patient linkage key was manually created as a benchmark to compare false positives and false negatives of the various candidate methods.
 ResultsIn general, adding data elements to the linkage methodology increased false negatives, decreased false positives and reduced the number of records that could be included in the linkage. Data quality issues affecting linkages varied across the data holdings and by jurisdiction.
 The final standard linkage methodology is simple and deterministic: link records by jurisdiction and HCN and exclude records where HCN is used by multiple people, for example, newborns who share their HCN with their mother.
 ConclusionArriving at a methodology that balanced the need for inclusiveness, simplicity and precision required collaboration, compromise, analysis and innovation. To date, the new client linkage standard has been implemented for ad-hoc analysis as well as in the data warehouse for our new web analytics tool. The implementation of the new client linkage standard represents an important foundational step towards creating an integrated analytics tool that spans across the continuum of care.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.008 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.004 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".