S841 Training Machine Learning Models for the Assessment of the Endoscopic Mayo Score in Ulcerative Colitis: A Systematic Review
Bibliographic record
Abstract
Introduction: The endoscopic Mayo Score (eMS) is utilized to provide an objective assessment of therapeutic endpoints in Ulcerative Colitis (UC) clinical trials, however inter- and intra-observer variability in assignment of eMS grades remains a challenge, even for trained central readers. Advancements in machine learning (ML) technology offer a potential solution to improve the consistency of endoscopic scoring. Methods employed in the development of these models may impact their accuracy in trials. The objective of this study is to provide a systematic review on training of ML models to generate an automated eMS grade on a full-length endoscopic procedure recording from patients with UC. Methods: Our review includes full-length manuscripts on ML video-level eMS prediction models from human studies in UC published in English. PubMed/MEDLINE, EMBASE, and Web of Science were systematically searched on December 31, 2023, and supplemented by reference checks and Google search. Three rounds of title screening were conducted. Information on ML models and training data were extracted independently by 2 authors, with disparities resolved through discussion. Results: Seven studies met criteria for inclusion. Five studies utilized endoscopic video recordings as the source data for model development (dataset size ranged from 134-1,881 videos), and 2 utilized still images captured from previous endoscopic procedures (both with a dataset size of 16,514 images). The final output of the model is an ordinal eMS grade (0, 1, 2, 3) in 6 studies, while one study generated a binary eMS grade in 3 ways (eMS >=1, eMS >=2, eMS >=3). Model architectures generally consisted of a few key components that contributed to the final video level output, including an informative image ML classifier (n=6), a still image eMS ML classifier (n=6), and video level eMS aggregation through statistical (n=5) or ML (n=2) techniques. Data labeling strategies to train still image and video level ML eMS classifiers varied across studies in the type of endoscopic data labeled, reading paradigm, and personnel. Conclusion: Several studies have trained ML models to assess the eMS in UC endoscopic videos. Training plays an important role in model generalizability, and variation in methodology may inform model performance. Awareness from clinicians, regulators, investigators, and industry on how these models are built will make them more adept in selecting the most appropriate models for future studies (see Figure 1, Table 1).Figure 1.: Overview of eMS labels applied to endoscopic data for model training. Labels are provided through a human annotation workflow with various paradigms described below. (A) eMS labels for training of image classification ML models (applied in 6 studies). (B) eMS labels for training of video classification ML models (applied in 2 studies). Studies in italics indicate the use of weak labels where a higher-level label is automatically assigned at a more granular level, such as automatically assigning an image a label based on the label at the video level. Table 1. - Description of training datasets and ML model characteristics for automated video-level eMS assessments in UC Study name and year Source of data Type of source endoscopic data Quantity of data Number of patients Final video-level output Summary of architecture Method for video-level aggregation Stidham et al, 2019 Routine care Images 16,514 images 3,082 Ordinal eMS grade (0, 1, 2, 3) 1. Still-image eMS classifier,2. Video-level eMS aggregation Proportion of frame eMS grades Becker et al, 2021 Clinical trial Videos 1,672 videos 1,105 Binary eMS grade, 3 ways: eMS >=1, eMS >=2, eMS >=3 1. Informative image classifier,2. Still-image eMS classifier,3. Video-level eMS aggregation Average of frame eMS grades Gottlieb et al, 2021 Clinical trial Videos 795 videos 249 Ordinal eMS grade (0, 1, 2, 3) 1. Informative image classifier,2. Still-image feature extractor,3. Video-level eMS aggregation RNN Schwab et al, 2021 Clinical trial Videos 1,881 videos 726 Ordinal eMS grade (0, 1, 2, 3) 1. Informative image classifier,2. Still-image eMS classifier,3. Video-level eMS aggregation Frame with the maximum eMS grade Yao et al, 2021 Routine care Images 16,514 images 3,082 Ordinal eMS grade (0, 1, 2, 3) 1. Informative image classifier,2. Still-image eMS classifier,3. Video-level eMS aggregation Proportion of frame eMS grades Byrne et al, 2023 Routine care Videos 134 videos Not specified Ordinal eMS grade (0, 1, 2, 3) 1. Informative image classifier,2. Still-image eMS classifier,3. Clip-level eMS assignment,4. Video-level eMS aggregation Clip with the maximum eMS grade (clip-level aggregation done via proportion of frame eMS grades) Rubin et al, 2023 Clinical trial Videos 793 videos 294 Ordinal eMS grade (0, 1, 2, 3) 1. Informative image classifier,2. Still-image feature and eMS classifier,3. Clip-level eMS assignment,4. Video-level eMS aggregation RNN (clip-level aggregation done via RNN) RNN, recurrent neural network.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.003 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".