MétaCan
Menu
← Retour à la cohorte
Enregistrement W4413051733 · doi:10.3233/shti250952

A Scoping Review of AI/ML Algorithm Updating Practices for Model Continuity and Patient Safety Using a Simplified Checklist

2025· review· en· W4413051733 sur OpenAlexaff
Ahmed Otokiti, Makuochukwu Maryann Ozoude, Ilse Siguachi, Leyla Warsame, Karmen S. Williams, Seyi John Akinloye

Notice bibliographique

RevueStudies in health technology and informatics · 2025
Typereview
Langueen
DomaineMedicine
ThématiqueArtificial Intelligence in Healthcare and Education
Établissements canadiensSouth Health Campus
Organismes subventionnairesnon disponible
Mots-clésChecklistMachine learningPsycINFOMEDLINESystematic reviewArtificial intelligenceCochrane LibraryComputer scienceReliability (semiconductor)MedicineData miningRandomized controlled trialPsychologyPathology

Résumé

récupéré en direct d'OpenAlex

The ubiquity of clinical artificial intelligence (AI) and machine learning (ML) models necessitates measures to ensure the reliability of model output over time. Previous reviews have highlighted the lack of external validation for most clinical models, but a comprehensive review assessing the current priority given to clinical model updating is lacking. The objective of this study was to analyze studies of clinical AI models based on PRISMA guidelines. Additionally, a new simple checklist/score system was developed and employed to screen the quality of published AI/ML models. The primary aim was to understand the extent to which clinical model updating is prioritized in current research. We conducted a systematic analysis of studies on clinical AI models, adhering to PRISMA guidelines. To assess the quality of the models, we introduced a new checklist/score and considered demographic composition based on ethnicity or race. This comprehensive approach aimed to provide a thorough evaluation of the current landscape of clinical AI models. A comprehensive literature search was conducted using Ovid Embase, Ovid MEDLINE, Ovid PsycINFO, Web of Science Core Collection, Scopus, and the Cochrane Library. Inclusion criteria encompassed AI and ML studies involving clinically predictive or prognostic modeling, human studies with algorithms, articles using supervised learning methods, articles using at least two predictor variables, and studies including randomized controlled trials, prospective and retrospective cohorts, case-control studies, and case-cohort studies. Studies that did not meet these inclusion criteria were excluded. This methodology ensures a thorough and systematic evaluation of clinical AI models. The results of our analysis revealed that only 9% of the reviewed 390 AI/ML studies on sampled models stated an intention or method to update their models in the future. 98% of the AI/ML models in our review were in the research phase, and only 2 % were in the production phase. Furthermore, a mere 12% reported following best practice standards for model development. Notably, 84% of the studies did not provide demographic composition based on ethnicity or race. These findings shed light on the characteristics of recent clinical models and underscore the prevalence of research phase models built on proprietary data, limiting independent verification and validation of model output. In conclusion, our review emphasizes the need for increased attention to the updating of clinical AI models, as a significant portion of studies currently lack commitment to future model updates. The low adherence to best practice standards for model development also highlights areas for improvement in the field. Furthermore, the absence of demographic information in a substantial number of studies raises concerns about the generalizability and equitable application of these models. These findings shed light on the characteristics of recent clinical models and underscore the prevalence of research phase models built on proprietary data, limiting independent verification and validation of model output is also a big concern for patient safety. Addressing these issues is crucial for advancing the reliability and inclusivity of clinical AI and ML applications.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,115
score de la tête « metaresearch » (Gemma)0,315
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche
Catégories consensuellesaucune
DomaineSignal candidat: Méthodes · Signal consensuel: aucune
Devis d'étudeSignal candidat: Revue systématique · Signal consensuel: Revue systématique
GenreSignal candidat: Synthèse · Signal consensuel: Synthèse
Score de désaccord entre enseignants0,885
Score d'incertitude au seuil0,610

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,1150,315
Méta-épidémiologie (sens strict)0,0040,002
Méta-épidémiologie (sens large)0,0080,014
Bibliométrie0,0480,037
Études des sciences et des technologies0,0030,003
Communication savante0,0080,010
Science ouverte0,0080,008
Intégrité de la recherche0,0050,004
Charge utile insuffisante (le modèle a refusé de juger)0,0090,002

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,321
Tête enseignante GPT0,582
Écart entre enseignants0,261 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Devis d'étudeRevue systématique
DomaineMéthodes
GenreSynthèse

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2025
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueStudies in health technology and informatics→Même sujetArtificial Intelligence in Healthcare and Education→Travaux en français237 207→