Real-world insights from an AI-ECG decision support platform: AI vs. blinded HCP STEMI ECG interpretation
Bibliographic record
Abstract
Abstract Background Early and accurate identification of ST-elevation myocardial infarction (STEMI) is critical for timely reperfusion therapy and improved patient outcomes. However, the accuracy of STEMI detection via ECG interpretation varies among healthcare professionals (HCPs). Artificial intelligence (AI)-assisted ECG interpretation has the potential to reduce this variability and enhance diagnostic accuracy, particularly in real-world settings with diverse provider expertise. Purpose To compare the diagnostic performance of AI-assisted ECG analysis and self-reported HCP interpretations in STEMI detection using a real-world dataset from an AI-integrated ECG analysis platform. Methods A total of 4,598 consecutive 12-lead ECGs were uploaded to an AI ECG platform between June 2023 and January 2024 by 1,423 users for AI decision support. Two expert physicians provided the reference standard, classifying 1,527 (33%) ECGs as STEMI. The AI ECG Model was evaluated on the full dataset. For 1,423 (30%) ECGs, 356 HCPs provided a self-declared initial interpretation, blinded to the AI output. These HCPs included 65 (18%) cardiologists, 199 (56%) non-cardiologist physicians, and 92 (26%) non-physician HCPs. Diagnostic performance for STEMI detection was assessed using sensitivity, specificity, PPV, NPV, and F1 score, with 95% confidence intervals calculated using the Wilson Score Interval and statistical significance determined via the Chi-Square Test. Results The AI ECG Model significantly outperformed HCPs across all evaluated metrics (p < 0.001). When tested on 4,598 ECGs, AI achieved 93.5% sensitivity, 87.0% specificity, 78.0% PPV, 96.4% NPV, and an F1 score of 85.1%. In contrast, HCPs across all training levels had 84.6% sensitivity, 73.2% specificity, 55.8% PPV, 92.3% NPV, and an F1 score of 67.2%. Performance comparison by HCP role showed that cardiologists had significantly higher PPV (66.9%) than non-cardiologist physicians (56.7%) and non-physicians (52.0%) (p = 0.012), though sensitivity, specificity, and NPV did not differ significantly. AI remained superior to all subgroups (p < 0.05). Conclusion The AI ECG Model demonstrated significantly higher diagnostic accuracy than healthcare professionals overall, including cardiologists. AI-assisted ECG analysis may improve early STEMI detection, reduce misdiagnoses, and support clinical decision-making. Further research is warranted to evaluate its impact on time to reperfusion and patient outcomes.Figure 1
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".