S841 Training Machine Learning Models for the Assessment of the Endoscopic Mayo Score in Ulcerative Colitis: A Systematic Review
Notice bibliographique
Résumé
Introduction: The endoscopic Mayo Score (eMS) is utilized to provide an objective assessment of therapeutic endpoints in Ulcerative Colitis (UC) clinical trials, however inter- and intra-observer variability in assignment of eMS grades remains a challenge, even for trained central readers. Advancements in machine learning (ML) technology offer a potential solution to improve the consistency of endoscopic scoring. Methods employed in the development of these models may impact their accuracy in trials. The objective of this study is to provide a systematic review on training of ML models to generate an automated eMS grade on a full-length endoscopic procedure recording from patients with UC. Methods: Our review includes full-length manuscripts on ML video-level eMS prediction models from human studies in UC published in English. PubMed/MEDLINE, EMBASE, and Web of Science were systematically searched on December 31, 2023, and supplemented by reference checks and Google search. Three rounds of title screening were conducted. Information on ML models and training data were extracted independently by 2 authors, with disparities resolved through discussion. Results: Seven studies met criteria for inclusion. Five studies utilized endoscopic video recordings as the source data for model development (dataset size ranged from 134-1,881 videos), and 2 utilized still images captured from previous endoscopic procedures (both with a dataset size of 16,514 images). The final output of the model is an ordinal eMS grade (0, 1, 2, 3) in 6 studies, while one study generated a binary eMS grade in 3 ways (eMS >=1, eMS >=2, eMS >=3). Model architectures generally consisted of a few key components that contributed to the final video level output, including an informative image ML classifier (n=6), a still image eMS ML classifier (n=6), and video level eMS aggregation through statistical (n=5) or ML (n=2) techniques. Data labeling strategies to train still image and video level ML eMS classifiers varied across studies in the type of endoscopic data labeled, reading paradigm, and personnel. Conclusion: Several studies have trained ML models to assess the eMS in UC endoscopic videos. Training plays an important role in model generalizability, and variation in methodology may inform model performance. Awareness from clinicians, regulators, investigators, and industry on how these models are built will make them more adept in selecting the most appropriate models for future studies (see Figure 1, Table 1).Figure 1.: Overview of eMS labels applied to endoscopic data for model training. Labels are provided through a human annotation workflow with various paradigms described below. (A) eMS labels for training of image classification ML models (applied in 6 studies). (B) eMS labels for training of video classification ML models (applied in 2 studies). Studies in italics indicate the use of weak labels where a higher-level label is automatically assigned at a more granular level, such as automatically assigning an image a label based on the label at the video level. Table 1. - Description of training datasets and ML model characteristics for automated video-level eMS assessments in UC Study name and year Source of data Type of source endoscopic data Quantity of data Number of patients Final video-level output Summary of architecture Method for video-level aggregation Stidham et al, 2019 Routine care Images 16,514 images 3,082 Ordinal eMS grade (0, 1, 2, 3) 1. Still-image eMS classifier,2. Video-level eMS aggregation Proportion of frame eMS grades Becker et al, 2021 Clinical trial Videos 1,672 videos 1,105 Binary eMS grade, 3 ways: eMS >=1, eMS >=2, eMS >=3 1. Informative image classifier,2. Still-image eMS classifier,3. Video-level eMS aggregation Average of frame eMS grades Gottlieb et al, 2021 Clinical trial Videos 795 videos 249 Ordinal eMS grade (0, 1, 2, 3) 1. Informative image classifier,2. Still-image feature extractor,3. Video-level eMS aggregation RNN Schwab et al, 2021 Clinical trial Videos 1,881 videos 726 Ordinal eMS grade (0, 1, 2, 3) 1. Informative image classifier,2. Still-image eMS classifier,3. Video-level eMS aggregation Frame with the maximum eMS grade Yao et al, 2021 Routine care Images 16,514 images 3,082 Ordinal eMS grade (0, 1, 2, 3) 1. Informative image classifier,2. Still-image eMS classifier,3. Video-level eMS aggregation Proportion of frame eMS grades Byrne et al, 2023 Routine care Videos 134 videos Not specified Ordinal eMS grade (0, 1, 2, 3) 1. Informative image classifier,2. Still-image eMS classifier,3. Clip-level eMS assignment,4. Video-level eMS aggregation Clip with the maximum eMS grade (clip-level aggregation done via proportion of frame eMS grades) Rubin et al, 2023 Clinical trial Videos 793 videos 294 Ordinal eMS grade (0, 1, 2, 3) 1. Informative image classifier,2. Still-image feature and eMS classifier,3. Clip-level eMS assignment,4. Video-level eMS aggregation RNN (clip-level aggregation done via RNN) RNN, recurrent neural network.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,016 | 0,071 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,001 |
| Méta-épidémiologie (sens large) | 0,008 | 0,010 |
| Bibliométrie | 0,008 | 0,007 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,003 | 0,003 |
| Science ouverte | 0,003 | 0,001 |
| Intégrité de la recherche | 0,002 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,007 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».