Developing EMR-based algorithms to Identify hospital adverse events for health system performance evaluation and improvement: Study protocol
Bibliographic record
Abstract
BACKGROUND: Measurement of care quality and safety mainly relies on abstracted administrative data. However, it is well studied that administrative data-based adverse event (AE) detection methods are suboptimal due to lack of clinical information. Electronic medical records (EMR) have been widely implemented and contain detailed and comprehensive information regarding all aspects of patient care, offering a valuable complement to administrative data. Harnessing the rich clinical data in EMRs offers a unique opportunity to improve detection, identify possible risk factors of AE and enhance surveillance. However, the methodological tools for detection of AEs within EMR need to be developed and validated. The objectives of this study are to develop EMR-based AE algorithms from hospital EMR data and assess AE algorithm's validity in Canadian EMR data. METHODS: Patient EMR structured and text data from acute care hospitals in Calgary, Alberta, Canada will be linked with discharge abstract data (DAD) between 2010 and 2020 (n~1.5 million). AE algorithms development. First, a comprehensive list of AEs will be generated through a systematic literature review and expert recommendations. Second, these AEs will be mapped to EMR free texts using Natural Language Processing (NLP) technologies. Finally, an expert panel will assess the clinical relevance of the developed NLP algorithms. AE algorithms validation: We will test the newly developed AE algorithms on 10,000 randomly selected EMRs between 2010 to 2020 from Calgary, Alberta. Trained reviewers will review the selected 10,000 EMR charts to identify AEs that had occurred during hospitalization. Performance indicators (e.g., sensitivity, specificity, positive predictive value, negative predictive value, F1 score, etc.) of the developed AE algorithms will be assessed using chart review data as the reference standard. DISCUSSION: The results of this project can be widely implemented in EMR based healthcare system to accurately and timely detect in-hospital AEs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.092 | 0.107 |
| Meta-epidemiology (narrow) | 0.004 | 0.003 |
| Meta-epidemiology (broad) | 0.005 | 0.006 |
| Bibliometrics | 0.005 | 0.005 |
| Science and technology studies | 0.004 | 0.003 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.005 | 0.003 |
| Research integrity | 0.004 | 0.005 |
| Insufficient payload (model declined to judge) | 0.048 | 0.011 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".