Customization of a Severity of Illness Score Using Local Electronic Medical Record Data
Bibliographic record
Abstract
PURPOSE: Severity of illness (SOI) scores are traditionally based on archival data collected from a wide range of clinical settings. Mortality prediction using SOI scores tends to underperform when applied to contemporary cases or those that differ from the case-mix of the original derivation cohorts. We investigated the use of local clinical data captured from hospital electronic medical records (EMRs) to improve the predictive performance of traditional severity of illness scoring. METHODS: We conducted a retrospective analysis using data from the Multiparameter Intelligent Monitoring in Intensive Care II (MIMIC-II) database, which contains clinical data from the Beth Israel Deaconess Medical Center in Boston, Massachusetts. A total of 17 490 intensive care unit (ICU) admissions with complete data were included, from 4 different service types: medical ICU, surgical ICU, coronary care unit, and cardiac surgery recovery unit. We developed customized SOI scores trained on data from each service type, using the clinical variables employed in the Simplified Acute Physiology Score (SAPS). In-hospital, 30-day, and 2-year mortality predictions were compared with those obtained from using the original SAPS using the area under the receiver-operating characteristics curve (AUROC) as well as the area under the precision-recall curve (AUPRC). Test performance in different cohorts stratified by severity of organ injury was also evaluated. RESULTS: Most customized scores (30 of 39) significantly outperformed SAPS with respect to both AUROC and AUPRC. Enhancements over SAPS were greatest for patients undergoing cardiovascular surgery and for prediction of 2-year mortality. CONCLUSIONS: Custom models based on ICU-specific data provided better mortality prediction than traditional SAPS scoring using the same predictor variables. Our local data approach demonstrates the value of electronic data capture in the ICU, of secondary uses of EMR data, and of local customization of SOI scoring.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.028 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".