P777 Deployment of an artificial intelligence tool for precision medicine in ulcerative colitis: Preliminary data from 8 globally distributed clinical sites
Bibliographic record
Abstract
Abstract Background Histological remission is an important target for Ulcerative Colitis (UC) treatment; however, scoring of histological images is time-consuming and prone to inter and intra-observer variability. Thus, a need exists for an accurate, reproducible, and reliable automated method. Previously, we demonstrated an Artificial Intelligence (AI) Tool using image processing and machine learning algorithms to measure histological disease activity using the Nancy index consistently and accurately.1 Here, we aim to enhance the capabilities of the AI Tool, by adding substantially more population-diversified training data while maintaining accuracy and robustness of results. Methods Eight global sites submitted 600 UC histological images. These were added to the 200 images previously used to train and validate the AI Tool. The 800-image dataset was divided into 2 groups: 90% used for training, 10% for testing. The novel AI algorithms were trained using state-of-the-art image processing and machine learning techniques based on deep learning and feature extraction. Cell and tissue regions of each training image were manually annotated, measured, and assigned a Nancy Index independently by 3 histopathologists, and used to further train the AI using over 43,000 characterisations. The AI Tool fully characterises histological images, identifying tissue types, cell types, cell numbers and locations, and automatically measures the Nancy Index for each image. Intra Class Correlation (ICC) and Confusion Matrix analyses were performed to evaluate the AI Tool and assess accuracy. Results The average ICC was 92.1% among the histopathologists and 91.1% between histopathologists and AI Tool, compared with 88.3% and 87.2% in the previous study.1 Confusion matrix analysis (Table 1) demonstrated the strongest correlation at the extremes of the Nancy Index, with 80% correlation between predicted and true labels for Nancy Scores of 0 or 4. When 2 adjacent scores were combined, correlations were stronger: 96% for a true Nancy score of 0 being predicted as 0 or 1, and 100% for a true Nancy score of 2 being predicted as 2 or 3. Conclusion By adding a larger number of images to the AI Tool training data, the robustness of the AI Tool was substantially improved while maintaining accuracy. The continued high correlation of AI Tool performance with the histopathologists reinforces the potential role for the AI Tool for IBD clinical applications. Fully characterising whole slides could standardise and validate an AI-driven scoring system for histology slides in IBD, eliminating the subjectivity of the human pathologist in assessment of disease activity. References: 1. Peyrin-Biroulet L, et al. J Crohn's and Colitis. 2022;16(Suppl 1):i105.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".