DOP13 Artificial Intelligence (AI) in endoscopy - Deep learning for detection and scoring of Ulcerative Colitis (UC) disease activity under multiple scoring systems
Bibliographic record
Abstract
Abstract Background Computer vision & deep learning(DL)to assess & help with tissue characterization of disease activity in Ulcerative Colitis(UC)through Mayo Endoscopic Subscore(MES)show good results in central reading for clinical trials.UCEIS(Ulcerative Colitis Endoscopic Index of Severity)being a granular index,may be more reflective of disease activity & more primed for artificial intelligence(AI). We set out to create UC detection & scoring,in a single tool & graphic user interface(GUI),improving accuracy & precision of MES & UCEIS scores & reducing the time elapsed between video collection,quality assurance & final scoring.We apply DL models to detect & filter scorable frames,assess quality of endoscopic recordings & predict MES & UCEIS scores in videos of patients with UC Methods We leveraged>375,000frames from endoscopy cases using Olympus scopes(190&180Series).Experienced endoscopists & 9 labellers tagged~22,000(6%)images showing normal, disease state(MES orUCEIS subscores)& non-scorable frames.We separate total frames in 3 categories:training(60%),testing(20%)&validation(20%).Using a Convolutional Neural Network(CNN)Inception V3,including a biopsy & post-biopsy detector,an out-of-the-body framework & blue light algorithm.Similar architecture for detection with multiple separate units & corresponding dense layers taking CNN to provide continuous scores for 5 separate outputs:MES,aggregate UCEIS & individual components Vascular Pattern,Bleeding & Ulcers. Results Multiple metrics evaluate detection models.Overall performance has an accuracy of~88% & a similar precision & recall for all classes. MAE(distance from ground truth)& mean bias(over/under-prediction tendency)are used to assess the performance of the scoring model.Our model performs well as predicted distributions are relatively close to the labelled,ground truth data & MAE & Bias for all frames are relatively low considering the magnitude of the scoring scale. To leverage all our models,we developed a practical tool that should be used to improve efficiency & accuracy of reading & scoring process for UC at different stages of the clinical journey. Conclusion We propose a DL approach based on labelled images to automate a workflow for improving & accelerating UC disease detection & scoring using MES & UCEIS scores. Our deep learning model shows relevant feature identification for scoring disease activity in UC patients, well aligned with both scoring guidelines,performance of experts & demonstrates strong promise for generalization.Going forward, we aim to continue developing our detection & scoring tool. With our detailed workflow supported by deep learning models, we have a driving function to create a precise & potentially superhuman level AI to score disease activity
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".