MétaCan
Menu
Back to cohort
Record W4306689887 · doi:10.1093/ecco-jcc/jjac152

Application of Deep Learning Models to Improve Ulcerative Colitis Endoscopic Disease Activity Scoring Under Multiple Scoring Systems

2022· article· en· W4306689887 on OpenAlexaff
Michael F. Byrne, Remo Panaccione, James E. East, Marietta Iacucci, Nasim Parsa, Rakesh Kalapala, D. Nageshwar Reddy, Hardik Rughwani, Aniruddha Pratap Singh, Sameer Berry, R Monsurate, Florian Soudan, Greta Laage, Enrico D Cremonese, L St-Denis, Paul Lemaître, Shima Nikfal, J Asselin, Milagros L Henkel, Simon Travis

Bibliographic record

VenueJournal of Crohn s and Colitis · 2022
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicInflammatory Bowel Disease
Canadian institutionsUniversity of CalgaryVancouver General HospitalUniversity of British Columbia
FundersNIHR Oxford Biomedical Research Centre
KeywordsArtificial intelligenceMedicineUlcerative colitisComputer scienceGround truthPreprocessorConvolutional neural networkDeep learningMachine learningDiseaseInternal medicine

Abstract

fetched live from OpenAlex

BACKGROUND AND AIMS: Lack of clinical validation and inter-observer variability are two limitations of endoscopic assessment and scoring of disease severity in patients with ulcerative colitis [UC]. We developed a deep learning [DL] model to improve, accelerate and automate UC detection, and predict the Mayo Endoscopic Subscore [MES] and the Ulcerative Colitis Endoscopic Index of Severity [UCEIS]. METHODS: A total of 134 prospective videos [1550 030 frames] were collected and those with poor quality were excluded. The frames were labelled by experts based on MES and UCEIS scores. The scored frames were used to create a preprocessing pipeline and train multiple convolutional neural networks [CNNs] with proprietary algorithms in order to filter, detect and assess all frames. These frames served as the input for the DL model, with the output being continuous scores for MES and UCEIS [and its components]. A graphical user interface was developed to support both labelling video sections and displaying the predicted disease severity assessment by the artificial intelligence from endoscopic recordings. RESULTS: Mean absolute error [MAE] and mean bias were used to evaluate the distance of the continuous model's predictions from ground truth, and its possible tendency to over/under-predict were excellent for MES and UCEIS. The quadratic weighted kappa used to compare the inter-rater agreement between experts' labels and the model's predictions showed strong agreement [0.87, 0.88 at frame-level, 0.88, 0.90 at section-level and 0.90, 0.78 at video-level, for MES and UCEIS, respectively]. CONCLUSIONS: We present the first fully automated tool that improves the accuracy of the MES and UCEIS, reduces the time between video collection and review, and improves subsequent quality assurance and scoring.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.917
Threshold uncertainty score0.516

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.012
GPT teacher head0.240
Teacher spread0.228 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations47
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Crohn s and ColitisSame topicInflammatory Bowel DiseaseFrench-language works237,207