Serial 12-Lead Electrocardiogram–Based Deep-Learning Model for Hospital Admission Prediction in Emergency Department Cardiac Presentations: Retrospective Cohort Study
Bibliographic record
Abstract
BACKGROUND: Emergency Department (ED) crowding is often attributed to a slow hospitalization process, leading to reduced quality of care. Predicting early disposition with cardiac-presenting patients is challenging: most are ultimately discharged, yet those with a cardiac etiology frequently require hospital admission. Existing scores rely on single-time-point data and often underperform when patient risk evolves during the visit. OBJECTIVE: To develop and validate a real-time deep-learning model that fuses serial 12-lead electrocardiogram (ECG) waveforms with sequential vitals and routinely available clinical data to predict hospital admission early in ED encounters. METHODS: We conducted a retrospective cohort study using the MIMIC-IV, MIMIC-IV-ED, and MIMIC-IV-ECG databases. Adults presenting with chest pain, dyspnea, syncope, or presyncope and at least one ECG within their ED stay were included. Two evaluation cohorts were defined: all stays with ≥1 ECG (N=30,421) and a subset with ≥2 ECGs during the encounter (N=11,273). To predict hospital admission, we first established two baseline models: a tabular model (random forest) trained on structured clinical variables including demographics, triage acuity, past medical history, medications, and laboratory results, and an ECG-only model that learned directly from raw 12-lead waveforms. We then developed a multimodal deep-learning model that combined ECGs with sequential vital signs as well as the same static tabular features. All models were restricted to data available during the stay up to the time of the last ECG. Performance was assessed with stratified 5-fold cross-validation using identical splits across models. RESULTS: The multimodal model achieved an Area Under Receiver Operating Characteristic (AUROC) of 0.911 when trained on all eligible stays. The model predicted disposition after the final ECG was taken, which was a median of 0.3 hours after triage and 4.6 hours before ED departure. Baseline models performed worse: the ECG-only model had an AUROC of 0.852, and the tabular random forest had an AUROC of 0.886. In the subset requiring at least two ECGs within the stay, ECG-only reached an AUROC of 0.859, and random forest, with the longer interval to chart tabular data, reached a higher AUROC of 0.911. The multimodal model had AUROC 0.924, and outperformed baselines in each cohort (paired DeLong P<.001). CONCLUSIONS: Serial ECGs, when integrated with evolving vitals and routine clinical features, enable accurate, early prediction of ED disposition in cardiac-presenting patients. This open-source, reproducible framework highlights the potential of multimodal deep learning to streamline ED flow, prioritize higher-risk cases, and detect evolving, time-critical pathology.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".