MétaCan
Menu
Back to cohort

Detecting and Remediating Harmful Data Shifts for the Responsible Deployment of Clinical AI Models

2025· article· en· W4411025380 on OpenAlexafffundabout
Vallijah Subasri, Amrit Krishnan, Ali Kore, Azra Dhalla, Deval Pandya, Bo Wang, David Malkin, Fahad Razak, Amol A. Verma, Anna Goldenberg, Elham Dolatabadi

Bibliographic record

VenueJAMA Network Open · 2025
Typearticle
Languageen
FieldMedicine
TopicArtificial Intelligence in Healthcare and Education
Canadian institutionsSt. Michael's HospitalUniversity Health NetworkUniversity of TorontoVector InstituteYork UniversityHospital for Sick Children
FundersCanadian Institutes of Health Research
KeywordsReceiver operating characteristicMedicineDemographicsHarmPipeline (software)Emergency medicineArtificial intelligenceMachine learningMedical emergencyComputer sciencePsychologyDemographyInternal medicine

Abstract

fetched live from OpenAlex

Importance: Clinical artificial intelligence (AI) systems are susceptible to performance degradation due to data shifts, which can lead to erroneous predictions and potential patient harm. Proactively detecting and mitigating these shifts is crucial for maintaining AI effectiveness and safety in clinical practice. Objectives: To develop and evaluate a proactive, label-agnostic monitoring pipeline to detect and mitigate harmful data shifts in clinical AI systems and to assess the use of transfer learning and continual learning strategies in maintaining model performance. Design, Setting, and Participants: This prognostic study was conducted using electronic health record data for admissions to general internal medicine wards of 7 large hospitals (5 academic and 2 community) in Toronto, Canada, between January 1, 2010, to August 31, 2020. Inpatients (aged ≥18 years) with a hospital stay of at least 24 hours were included. Data analysis was performed from January to August 2022. Exposures: Data shifts due to changes in hospital type, critical laboratory assays, patient demographics, admission type, and the COVID-19 pandemic. Main Outcomes and Measures: The primary outcome was predictive performance for all-cause in-hospital mortality within the next 2 weeks, evaluated using the area under the receiver operating characteristic curve (AUROC) and the area under the precision-recall curve (AUPRC). Data shifts were detected using a label-agnostic monitoring pipeline employing a black box shift estimator with maximum mean discrepancy testing. Results: Data were available for 143 049 adult inpatients (mean [SD] age, 67.8 [19.6] years; 50.7% female). Significant data shifts were detected as a result of changes in younger age groups and admissions from nursing homes and acute care centers, transferring from community to academic hospitals, and changes in brain natriuretic peptide and D-dimer. Transfer learning improved model performance of community hospitals in a hospital type-dependent manner (Delta AUROC [SD], 0.05 [0.03]; Delta AUPRC [SD], 0.06 [0.04]). During the COVID-19 pandemic, drift-triggered continual learning improved overall model performance (Delta AUROC [SD], 0.44 [0.02]; P = .007, Mann-Whitney U test). Conclusions and Relevance: In this prognostic study, a proactive, label-agnostic monitoring pipeline detected harmful data shifts for a clinical AI system predicting in-hospital mortality. Transfer learning and drift-triggered continual learning strategies mitigated performance degradation, maintaining model performance across health care settings. These findings suggest that the approach used here may ensure the robust and equitable deployment of clinical AI models. Future research should explore the generalizability of this framework across diverse clinical domains, data modalities, and longer deployment periods to further validate its effectiveness.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.006
metaresearch head score (Gemma)0.003
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.892
Threshold uncertainty score0.416

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0060.003
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.545
GPT teacher head0.572
Teacher spread0.027 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations43
Published2025
Admission routes3
Has abstractyes

Explore more

Same venueJAMA Network OpenSame topicArtificial Intelligence in Healthcare and EducationFrench-language works237,207