MétaCan
Menu
Back to cohort
Record W4413252984 · doi:10.1136/bmjdhai-2025-000114

Development and prospective evaluation of a machine learning model to predict serious cardiac outcomes among paediatric cardiac inpatients

2025· article· en· W4413252984 on OpenAlexaff
Santiago Eduardo Arciniegas, Adam P. Yan, Adam Rapoport, Aamir Jeewa, Rugambwa Michael Muhame, Lin Guo, Agata Wolochacz, Karim Jessa, George Tomlinson, C Ferracci, Lillian Sung, Anne I. Dipchand, Kate Nelson

Bibliographic record

VenueBMJ Digital Health & AI · 2025
Typearticle
Languageen
FieldComputer Science
TopicMachine Learning in Healthcare
Canadian institutionsCanadian Patient Safety InstituteSickKids FoundationToronto General HospitalCasey HouseInstitute for Clinical Evaluative SciencesHospital for Sick Children
Fundersnot available
KeywordsRetrospective cohort studyMedicineReceiver operating characteristicLogistic regressionProspective cohort studyEmergency medicineInternal medicine

Abstract

fetched live from OpenAlex

Objectives Objectives were to develop a machine learning (ML) model based on electronic health record data to predict the risk of a serious cardiac outcome within the next 3 months among patients admitted to the cardiology service using retrospective data, and to evaluate the model prospectively in a silent trial (predictions not provided to clinicians). Methods and analysis Admissions between 2 June 2018 to 21 August 2023 (retrospective) and 10 May 2024 to 26 October 2024 (prospective) to the cardiology service were included. Data were a curated and validated source named SickKids Enterprise-wide Data in Azure Repository. Prediction time was the morning following admission. The label was a composite outcome consisting of ventricular assist device procedure, heart transplant waitlisting or death within 3 months. We trained models using L2-regularised logistic regression, LightGBM and XGBoost. Training cohorts include the target cohort and all inpatient admissions. Results The best-performing model in the retrospective phase was LightGBM trained on all inpatients. There were 51 571 admissions used for model development in the retrospective phase and 515 admissions in the prospective silent trial. The number of features in the final model was 7553. The area under the receiver operating characteristic curve was 0.88 (95% CI 0.88 to 0.89) for retrospective and 0.82 (95% CI 0.79 to 0.83) for prospective silent trial phases. Based on a threshold selected during the retrospective phase, silent trial positive and negative predictive values were 0.19 and 0.97, respectively. Conclusions We created an ML model to predict serious cardiac outcomes using a deployment-aware framework leveraging real-world data. Postdeployment evaluation will be an important future goal.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.346
Threshold uncertainty score0.959

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0020.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.001
Science and technology studies0.0000.000
Scholarly communication0.0000.001
Open science0.0000.001
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.021
GPT teacher head0.343
Teacher spread0.323 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueBMJ Digital Health & AISame topicMachine Learning in HealthcareFrench-language works237,207