MétaCan
Menu
Back to cohort
Record W4406703674 · doi:10.1093/ecco-jcc/jjae190.0625

P0451 Using a machine learning model to assess agreement between inflammation in the rectosigmoid and entire colon in patients with ulcerative colitis

2025· article· en· W4406703674 on OpenAlexaboutno aff
Alexander George, Klaus Gottlieb, Shrujal S. Baxi, William Eastman, D Colucci, Chakib Battioui, Y Wang, Josh Lehrer, Pavel Brodskiy, Mohammad Haft‐Javaherian, D T Rubin

Bibliographic record

VenueJournal of Crohn s and Colitis · 2025
Typearticle
Languageen
FieldMedicine
TopicRadiomics and Machine Learning in Medical Imaging
Canadian institutionsnot available
Fundersnot available
KeywordsMedicineUlcerative colitisGastroenterologyInflammationColitisInternal medicineInflammatory bowel diseaseDisease

Abstract

fetched live from OpenAlex

Abstract Background Limited endoscopy of the rectosigmoid colon is sometimes used to assess endpoints in clinical trials for ulcerative colitis (UC). However, there is inherent inter- and intra-rater variability in endoscopic assessment of inflammation, and regulatory agencies have suggested that full colonoscopy is preferred1,2,3. We used a machine learning (ML) model to compare the degree of inflammation of full-length endoscopy videos to that seen in the distal colon and rectum. Methods We used a previously developed ML model that predicts the endoscopy subscore4, a component of the modified Mayo Score, and applied it to 49 full-length endoscopy videos randomly selected and stratified by endoscopic severity from the Phase 3 induction trial for mirikizumab in UC (NCT03518086). Each video was preprocessed by a human reviewer to identify the point of maximal extent and the end of the procedure. The withdrawal portion of the videos were divided into 15, 30, or 60 second clips (a series of smaller segments of video), and the ML model generated an endoscopy subscore for each clip (clip-level endoscopy subscore) and for the entire video (video-level endoscopy subscore). We assigned the final two clips during withdrawal as an assessment of the sigmoid and rectum. Therefore, in this analysis, we compared the ML endoscopy subscore grade of the final two clips to the video-level endoscopy subscore. Kappa statistics of variability between scores were performed. Results There was substantial agreement between the video-level endoscopy subscore and that of the last 2 clips (Table 1). For 60-second clips, the agreement rate between the video-level endoscopy subscore and clip 1 was 0.67 and that with clip 2 was 0.76. When comparing the video-level endoscopy subscore to clip 1 or clip 2, the agreement rate was 0.86, and this was greater than the agreement rate when looking at the higher grade of clip 1 or clip 2 (0.73). The 60 second clips had greater agreement rates than 30 second or 15 second clips (Table 1 and 2). The endoscopy subscores of clip 1 or clip 2 had excellent agreement with video-level grades 0-3 (Table 2), and this was best observed with 60 second clips (Table 2A, kappa 0.944). Conclusion The ML model demonstrated excellent agreement between clip-level endoscopy subscore assessments of the distal most portion of the withdrawal videos compared to the video-level endoscopy subscore, and this was best observed with 60 second clips. These findings support the use of distal colon and rectum assessment to determine the grade degree of endoscopic inflammation of the full colonoscopy. We propose further prospective study of the use of this ML model in this setting. References Feagan, B. G., Khanna, R., Sandborn, W. J., Vermeire, S., Reinisch, W., Su, C., . . . & Sands, B. E. (2021). Agreement between local and central reading of endoscopic disease activity in ulcerative colitis: Results from the tofacitinib OCTAVE trials. Alimentary Pharmacology & Therapeutics, 54(11-12), 1442-1453. Hashash, J. G., Yu Ci Ng, F., Farraye, F. A., Wang, Y., Colucci, D. R., Baxi, S., . . . & Melmed, G. Y. (2024). Inter-and intraobserver variability on endoscopic scoring systems in crohn’s disease and ulcerative colitis: A systematic review and meta-analysis. Inflammatory Bowel Diseases, izae051. Food and Drug Administration. Ulcerative Colitis: Developing Drugs for Treatment [Internet]. 2022 [cited 2024 Oct 17]. Available from: https://www.fda.gov/regulatory-information/search-fda-guidance-documents/ulcerative-colitis-developing-drugs-treatment Rubin, D. T., Gottlieb, K., Colombel, J. F., Schott, J. P., Erisson, L., Prucka, B., . . . & McGill, J. (2023). Development of a novel ulcerative colitis Endoscopic Mayo Score prediction model using machine learning. Gastro Hep Advances, 2(7), 935-942.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.019
metaresearch head score (Gemma)0.044
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.019
Threshold uncertainty score0.098

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0190.044
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0020.001
Science and technology studies0.0000.001
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0030.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.018
GPT teacher head0.303
Teacher spread0.285 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Crohn s and ColitisSame topicRadiomics and Machine Learning in Medical ImagingFrench-language works237,207