MétaCan
Menu
Back to cohort
Record W4318577082 · doi:10.1093/ecco-jcc/jjac190.0907

P777 Deployment of an artificial intelligence tool for precision medicine in ulcerative colitis: Preliminary data from 8 globally distributed clinical sites

2023· article· en· W4318577082 on OpenAlexaff
Laurent Peyrin‐Biroulet, Daniel S. Rubin, Christopher R. Weber, Shashi Adsul, Marina Lopes Freitas de Freire, Luc Biedermann, Viktor H. Koelzer, Brian Bressler, Jan Hendrik Niess, Madlaina Matter, Uri Kopylov, Iris Barshack, Fernando Magro, Fátima Carneiro, N Maharshak, A Greenberg, Sara Hart, Jashmid Dehmeshki, Olga Kubassova

Bibliographic record

VenueJournal of Crohn s and Colitis · 2023
Typearticle
Languageen
FieldMedicine
TopicRadiomics and Machine Learning in Medical Imaging
Canadian institutionsUniversity of British Columbia
Fundersnot available
KeywordsArtificial intelligenceConfusion matrixComputer scienceMachine learningPopulationConvolutional neural networkAutomated methodRobustness (evolution)ConfusionPattern recognition (psychology)Medicine

Abstract

fetched live from OpenAlex

Abstract Background Histological remission is an important target for Ulcerative Colitis (UC) treatment; however, scoring of histological images is time-consuming and prone to inter and intra-observer variability. Thus, a need exists for an accurate, reproducible, and reliable automated method. Previously, we demonstrated an Artificial Intelligence (AI) Tool using image processing and machine learning algorithms to measure histological disease activity using the Nancy index consistently and accurately.1 Here, we aim to enhance the capabilities of the AI Tool, by adding substantially more population-diversified training data while maintaining accuracy and robustness of results. Methods Eight global sites submitted 600 UC histological images. These were added to the 200 images previously used to train and validate the AI Tool. The 800-image dataset was divided into 2 groups: 90% used for training, 10% for testing. The novel AI algorithms were trained using state-of-the-art image processing and machine learning techniques based on deep learning and feature extraction. Cell and tissue regions of each training image were manually annotated, measured, and assigned a Nancy Index independently by 3 histopathologists, and used to further train the AI using over 43,000 characterisations. The AI Tool fully characterises histological images, identifying tissue types, cell types, cell numbers and locations, and automatically measures the Nancy Index for each image. Intra Class Correlation (ICC) and Confusion Matrix analyses were performed to evaluate the AI Tool and assess accuracy. Results The average ICC was 92.1% among the histopathologists and 91.1% between histopathologists and AI Tool, compared with 88.3% and 87.2% in the previous study.1 Confusion matrix analysis (Table 1) demonstrated the strongest correlation at the extremes of the Nancy Index, with 80% correlation between predicted and true labels for Nancy Scores of 0 or 4. When 2 adjacent scores were combined, correlations were stronger: 96% for a true Nancy score of 0 being predicted as 0 or 1, and 100% for a true Nancy score of 2 being predicted as 2 or 3. Conclusion By adding a larger number of images to the AI Tool training data, the robustness of the AI Tool was substantially improved while maintaining accuracy. The continued high correlation of AI Tool performance with the histopathologists reinforces the potential role for the AI Tool for IBD clinical applications. Fully characterising whole slides could standardise and validate an AI-driven scoring system for histology slides in IBD, eliminating the subjectivity of the human pathologist in assessment of disease activity. References: 1. Peyrin-Biroulet L, et al. J Crohn's and Colitis. 2022;16(Suppl 1):i105.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.007
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.004
Threshold uncertainty score0.021

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0040.007
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.001
Science and technology studies0.0000.001
Scholarly communication0.0010.001
Open science0.0000.001
Research integrity0.0010.000
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.074
GPT teacher head0.407
Teacher spread0.334 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Crohn s and ColitisSame topicRadiomics and Machine Learning in Medical ImagingFrench-language works237,207