MétaCan
Menu
Back to cohort
Record W4413779804 · doi:10.14309/ajg.0000000000003756

Real-Life Clinical Validation of Artificial Intelligence-Assisted Detection and Differentiation of Pleomorphic Lesions in Capsule Endoscopy

2025· article· en· W4413779804 on OpenAlexaff
Miguel Mascarenhas, João Ferreira, João Afonso, Francisco Mendes, William Sonnier, Bruno Rosa, Tiago Ribeiro, Tiago Cúrdia Gonçalves, Miguel Martins, Pedro Emartino Bezerra Campelo, C Macedo, Pedro Cardoso, Joana Mota, Maria João Almeida, António Miguel Martins Pinto da Costa, Ana Pérez-Gonzalez, Jorge Mendoza, Erika Borges Fortes, Matheus Ferreira de Carvalho, Marcos Eduardo Lera dos Santos, Patrícia Andrade, Hélder Cardoso, Eduardo Horneaux de Moura, Cecílio Santander, J. A. Palma, José Cotter, Guilherme Macedo

Bibliographic record

VenueThe American Journal of Gastroenterology · 2025
Typearticle
Languageen
FieldMedicine
TopicColorectal Cancer Screening and Detection
Canadian institutionsArtificial Intelligence in Medicine (Canada)
Fundersnot available
KeywordsMedicineCapsule endoscopyCapsuleEndoscopyPathologyRadiology

Abstract

fetched live from OpenAlex

INTRODUCTION: Capsule endoscopy (CapE) is a minimally invasive procedure designed for small bowels' evaluation. However, prolonged reading times and a risk of missing clinically significant findings limit its potential. Prospective clinical validation studies of artificial intelligence (AI) for CapE remain scarce. Furthermore, existing studies focus on lesion detection, without addressing lesion differentiation. METHODS: The aim of a multicenter prospective validation study was to compare AI-assisted reading with conventional CapE reading. Three hundred thirty CapE videos from 3 devices across 7 centers and 4 countries were included. After conventional reading reports, AI-assisted reading was performed by an independent expert using a deep learning model to detect and differentiate pleomorphic small bowel lesions. Both reports were reviewed by an expert from an independent center, which decided in discrepant cases. AI-assisted and standard readings were evaluated through their accuracy, sensitivity, specificity, positive and negative predictive value, and small bowel lesion detection rate. RESULTS: AI-assisted reading detected 605 of 635 lesions identified by expert-based consensus, whereas standard reading identified 354 lesions. AI-assisted reading outperformed standard reading, with small bowel lesion detection rate of 96.1% vs 76.3% and sensitivity of 97.5% vs 78.2%. AI-assisted reading had a mean examination reading time of 203 seconds per examination. DISCUSSION: This was the first multicentric study proving AI-assisted CapE reading superiority compared with conventional reading. The inclusion of videos from multiple devices addresses the interoperability challenge, whereas including patients from 4 countries and 2 different continents assures a diverse demographic context. AI achieved gastroenterologist-level identification of small-bowel lesions, surpassing conventional reading methods in both lesion detection and characterization.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.014
metaresearch head score (Gemma)0.049
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.014
Threshold uncertainty score0.074

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0140.049
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0000.001
Bibliometrics0.0010.001
Science and technology studies0.0000.001
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0010.000
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.041
GPT teacher head0.342
Teacher spread0.300 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueThe American Journal of GastroenterologySame topicColorectal Cancer Screening and DetectionFrench-language works237,207