MétaCan
Menu
Back to cohort
Record W2891601420 · doi:10.23889/ijpds.v3i4.984

Pan-Canadian Real-World Health Data Network: Building a National Data Platform

2018· article· en· W2891601420 on OpenAlexaffabout
Mark Smith, Kim McGrail, Michael J. Schull, Alan Katz, Ted McDonald, P. Alison Paprica, J. Charles Victor, Lisa M. Lix, Dan Château, Brent Diverty

Bibliographic record

VenueInternational Journal for Population Data Science · 2018
Typearticle
Languageen
FieldDecision Sciences
Topicdemographic modeling and climate adaptation
Canadian institutionsCanadian Institute for Health InformationUniversity of New BrunswickUniversity of ManitobaInstitute for Clinical Evaluative SciencesUniversity of British ColumbiaVector InstituteManitoba Health
Fundersnot available
KeywordsComputer scienceData scienceWork (physics)Construct (python library)PopulationSocial network analysisWorld Wide WebMedicineEngineeringEnvironmental health

Abstract

fetched live from OpenAlex

IntroductionResearchers and decision makers from across Canada use linked provincial administrative data for analysis and to address research and policy questions. Currently there are several impediments to working harmoniously across provincial boundaries. A group of academic and policy researchers are working to address these multi-jurisdictional obstacles. Objectives and ApproachResearchers and data organizations from across Canada are working together as the Pan-Canadian Real-World Health Data Network PRHDN). PRHDN aims to: (1) create harmonized data, algorithms and analytic protocols, and (2) link administrative databases to other types of data, including electronic medical records, clinical trials records, “omics data” and records from pan-Canadian cohort studies. PRHDN’s vision is to construct a unified, documented infrastructure to advance pan-Canadian population-based research and analysis. This presentation incorporates material that is part of PRHDN’s response to a funding call to create national, collaborative infrastructure. ResultsScientists and staff at PRHDN organizations will create three main categories of infrastructure: 1) Algorithms: Reusable processes, ideally in the form of documented code, which implement a common approach or definition, e.g. to define cases or to create derived variables; 2) Harmonized Common Data: Based on the Sentinel model, we will establish a standardized subset of harmonized common data that are analysis-ready; 3) Common Analytic Protocols: Complementing work of the Canadian Network for Observational Drug Effect Studies (CNODES), we will establish processes for distributed analysis with common analytic protocols and meta-analysis of results to provide pan-Canadian estimates. Source data would remain within jurisdictional boundaries and only aggregate results would be pooled across jurisdictions. Details of these approaches will be presented. Conclusion/ImplicationsThis initiative will improve coordinated access to distributed data from across Canada that is built once then used by many stakeholders for a variety of purposes including: research, benchmarking, performance monitoring to identify gaps and opportunities for improvement, multi-jurisdictional evaluations of novel interventions and inter-jurisdictional comparisons.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.074
metaresearch head score (Gemma)0.112
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesScholarly communication, Open science
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.990
Threshold uncertainty score0.756

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0740.112
Meta-epidemiology (narrow)0.0010.002
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0160.024
Science and technology studies0.0120.005
Scholarly communication0.0190.013
Open science0.0100.020
Research integrity0.0030.004
Insufficient payload (model declined to judge)0.0140.007

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.527
GPT teacher head0.550
Teacher spread0.023 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2018
Admission routes2
Has abstractyes

Explore more

Same venueInternational Journal for Population Data ScienceSame topicdemographic modeling and climate adaptationFrench-language works237,207