A32 ARTIFICIAL INTELLIGENCE USE IN DIAGNOSIS & MONITORING OF INFLAMMATORY BOWEL DISEASE: A SCOPING REVIEW
Bibliographic record
Abstract
Abstract Background Inflammatory bowel diseases (IBD) are a family of immune-mediated conditions, which are increasing in incidence and prevalence worldwide. Assessment of IBD is done through endoscopy, video capsule endoscopy (VCE), histology, and various imaging modalities including ultrasound (US), computed tomography (CT), and magnetic resonance imaging (MRI). Considering the increasing complexities in the assessment of IBD, artificial intelligence (AI) is an important adjunct with potential to enhance diagnosis, drug response and prediction of disease course Aims We conducted a scoping review to assess AI in diagnosis, monitoring, and prognostication of patients with IBD, to aid in identification of gaps in knowledge to guide future research endeavors. Methods The scoping review protocol was adapted from the recommendations laid out by the Preferred Reporting Items for Systematic Reviews and Meta-Analysis - Scoping Review Extension (PRISMA-ScR). Electronic databases used in the literature search included MEDLINE, EMBASE, the Cochrane Library, Cumulative Index to Nursing and Allied Health Literature, and Engineering Village. Two reviewers independently screened the abstracts and titles first before performing full text review. A third review resolved any conflict where needed. All study types were included, and data extraction utilized Covidence. Studies were categorized based on the assessment modality, then themes including, diagnosis, grading activity, prognosis, and monitoring. Results A total of 140 studies were included in the final scoping review. The largest number of studies involved endoscopy at 72 (51%) citations, followed by VCE, histology, MRI, CT, and US at 30 (21%), 18 (13%), 13 (9%), 6 (4%), and 1 (0.7%) citation(s), respectively. When looking at themes, most endoscopy studies examined disease activity (65%) while diagnosis was the most common theme in VCE, MRI and CT (77%, 69% and 83%, respectively). Histologic studies focused on prognosis (89%) and the single US study evaluated both diagnosis and prognosis concomitantly. Amongst all the investigative modalities examined, monitoring of IBD was the least studied theme. Peak performance of AI models for grading disease activity during endoscopy was 98.7% compared to human clinicians with less variability observed. Conclusions With IBD diagnosis and assessment becoming increasingly complex, AI may be a useful adjunctive tool across multiple modalities. Evaluation of use of AI in US is lacking, despite gaining interest in non-invasive assessment of IBD. Further studies are needed incorporating AI use in US while also investigating its role in monitoring IBD disease activity. We hope this scoping review will serve as a future direction for subsequent research in this area. Funding Agencies:
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.021 | 0.088 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.006 | 0.009 |
| Bibliometrics | 0.030 | 0.026 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.006 | 0.006 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.004 | 0.002 |
| Insufficient payload (model declined to judge) | 0.011 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".