Current use and costs of electronic health records for clinical trial research: a descriptive study
Bibliographic record
Abstract
BACKGROUND: Electronic health records (EHRs) may support randomized controlled trials (RCTs). We aimed to describe the current use and costs of EHRs in RCTs, with a focus on recruitment and outcome assessment. METHODS: This descriptive study was based on a PubMed search of RCTs published since 2000 that evaluated any medical intervention with the use of EHRs. Cost information was obtained from RCT investigators who used EHR infrastructures for recruitment or outcome measurement but did not explore EHR technology itself. RESULTS: We identified 189 RCTs, most of which (153 [81.0%]) were carried out in North America and were published recently (median year 2012 [interquartile range 2009-2014]). Seventeen RCTs (9.0%) involving a median of 732 (interquartile range 73-2513) patients explored interventions not related to EHRs, including quality improvement, screening programs, and collaborative care and disease management interventions. In these trials, EHRs were used for recruitment (14 [82%]) and outcome measurement (15 [88%]). Overall, in most of the trials (158 [83.6%]), the outcome (including many of the most patient-relevant clinical outcomes, from unscheduled hospital admission to death) was measured with the use of EHRs. The per-patient cost in the 17 EHR-supported trials varied from US$44 to US$2000, and total RCT costs from US$67 750 to US$5 026 000. In the remaining 172 RCTs (91.0%), EHRs were used as a modality of intervention. INTERPRETATION: Randomized controlled trials are frequently and increasingly conducted with the use of EHRs, but mainly as part of the intervention. In some trials, EHRs were used successfully to support recruitment and outcome assessment. Costs may be reduced once the data infrastructure is established.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.097 | 0.396 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.004 |
| Bibliometrics | 0.024 | 0.034 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.007 | 0.009 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".