Real-Life Clinical Validation of Artificial Intelligence-Assisted Detection and Differentiation of Pleomorphic Lesions in Capsule Endoscopy
Bibliographic record
Abstract
INTRODUCTION: Capsule endoscopy (CapE) is a minimally invasive procedure designed for small bowels' evaluation. However, prolonged reading times and a risk of missing clinically significant findings limit its potential. Prospective clinical validation studies of artificial intelligence (AI) for CapE remain scarce. Furthermore, existing studies focus on lesion detection, without addressing lesion differentiation. METHODS: The aim of a multicenter prospective validation study was to compare AI-assisted reading with conventional CapE reading. Three hundred thirty CapE videos from 3 devices across 7 centers and 4 countries were included. After conventional reading reports, AI-assisted reading was performed by an independent expert using a deep learning model to detect and differentiate pleomorphic small bowel lesions. Both reports were reviewed by an expert from an independent center, which decided in discrepant cases. AI-assisted and standard readings were evaluated through their accuracy, sensitivity, specificity, positive and negative predictive value, and small bowel lesion detection rate. RESULTS: AI-assisted reading detected 605 of 635 lesions identified by expert-based consensus, whereas standard reading identified 354 lesions. AI-assisted reading outperformed standard reading, with small bowel lesion detection rate of 96.1% vs 76.3% and sensitivity of 97.5% vs 78.2%. AI-assisted reading had a mean examination reading time of 203 seconds per examination. DISCUSSION: This was the first multicentric study proving AI-assisted CapE reading superiority compared with conventional reading. The inclusion of videos from multiple devices addresses the interoperability challenge, whereas including patients from 4 countries and 2 different continents assures a diverse demographic context. AI achieved gastroenterologist-level identification of small-bowel lesions, surpassing conventional reading methods in both lesion detection and characterization.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.049 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".