Use of the CANHEART ‘big data’ registry to conduct a large randomized registry clinical trial to improve lipid management in Ontario, Canada
Bibliographic record
Abstract
IntroductionCreation of registries using linked population-based databases could potentially be used to conduct large registry-based clinical trials. As part of the CANHEART-Strategy for Patient-Oriented Research (SPOR) innovative clinical trials initiative, we explored the practicality of using the linked CANHEART registry to conduct a cluster-randomized trial aimed at improving lipid management.
 Objectives and ApproachThe CANHEART registry (www.canheart.ca) was created through the linkage of 19 population-based health databases in Ontario, Canada, providing individual-level socio-demographic, geographic, hospitalization, disease testing/screening, mortality, prescription medication, and behavior/lifestyle information. Using CANHEART defined eligibility criterion, small and medium-sized, high cardiovascular-risk health regions (defined as having acute myocardial infarction, stroke or cardiovascular death rates greater than the Ontario average) are being randomly allocated to receive either the intervention (availability of a lipid management ‘toolbox’) or standard care. Cohort linkages to additional years of data will occur regularly over the 3-year trial to ascertain the primary outcome of appropriate statin prescribing rates.
 
 \section*{Results}
 Record linkage enabled us to determine baseline characteristics of 835,345 patients aged 40-75 as of January 2016, being treated in the 28 study-eligible regions by 2,012 family physicians. Preceding the study, the baseline statin use rate was 35.7\% (in 66-75 year olds) across these regions and the cardiovascular event rate ranged from 3.78-5.64 events/1000 person-years. A randomization procedure yielded 14 regions in both the intervention and control arms which did not differ significantly in socio-demographic characteristics, traditional cardiovascular risk factors, disease history, prevalence of statin use, or access to healthcare indicators. Working groups have been established to operationalize the lipid management tools that will be made available in the intervention regions. Analysis of newly linked participant data will permit outcome ascertainment at trial completion.
 
 \section*{Conclusion/Implications}
 Our work demonstrates the feasibility of using the CANHEART ‘big data’ registry to conduct a large, cluster-randomized clinical trial aimed at improving lipid management, without requiring any primary data collection. Broader use of this methodology has the potential to change the existing paradigm for conducting pragmatic clinical trial research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".