MétaCan
Menu
Retour à la cohorte
Enregistrement W4414763854 · doi:10.1055/a-2706-0201

Artificial intelligence for colorectal polyp sizing: clinical promise and methodological challenges

2025· article· en· W4414763854 sur OpenAlexaff
Daniel von Renteln

Notice bibliographique

RevueEndoscopy · 2025
Typearticle
Langueen
DomaineMedicine
ThématiqueColorectal Cancer Screening and Detection
Établissements canadiensUniversité de Montréal
Organismes subventionnairesnon disponible
Mots-clésColonoscopyColorectal PolypMEDLINEApplications of artificial intelligenceColonic disease

Résumé

récupéré en direct d'OpenAlex

10.1055/a-2695-1978 Accurate determination of colorectal polyp size remains one of the most critical yet technically elusive metrics in endoscopy. In the current issue of Endoscopy , Antonelli et al. present an important and timely advance for colonoscopy practice: the first clinical, real-life evaluation of a novel artificial intelligence (AI)-based polyp sizing model [ 1 ]. This study deserves recognition not only for its novelty but also for its clinical relevance. Polyp size is a critical diagnostic variable that governs surveillance intervals, guides resection strategies, and indicates the probability of a lesion having advanced neoplasia. The authors demonstrate that AI-assisted polyp sizing solutions can meanwhile be prospectively tested during real-world procedures. Antonelli et al. move the field forward by showing that comprehensive computational support in everyday practice – combining detection, histology prediction, and now polyp sizing – is becoming a reality. The key difficulty in developing AI for polyp sizing lies in the definition of “ground truth” [ 2 ]. Visual size estimation, although ubiquitous, has repeatedly been shown to be inaccurate and systematically biased [ 3 ] [ 4 ]. Non-calibrated instruments, such as biopsy forceps or open snares, have been used as reference, yet these tools themselves lack precision and cannot be regarded as definitive standards for training or validation of AI sizing systems [ 2 ] [ 3 ] [ 5 ] [ 6 ]. Histopathology has traditionally served as a reference standard, but its limitations are now well recognized: formalin fixation induces tissue shrinkage of approximately 20%, while orientation and sectioning artifacts may further distort measurements [ 7 ] [ 8 ]. Consequently, studies based on non-calibrated tools or histopathological size carry an intrinsic risk of error that complicates downstream AI training and validation. The reliance on flawed surrogates underscores the need for more robust ground truth approaches. “Build and validated on reliable ground truth, however, AI will deliver the needed integrative solutions going beyond detection and histology prediction, to include accurate AI based sizing for colorectal polyps.” Several groups have meanwhile reported convolutional neural network (CNN)-based models for automated polyp sizing, highlighting both the potential and the pitfalls of this approach. While promising accuracy has been achieved under experimental conditions, attention has also been brought to the underappreciated technical complexities of endoscopic imaging. Lens characteristics – especially the fisheye’s field of view distortion – alter the apparent dimensions of polyps. Without proper ground truth and explicit accounting for optical distortions, AI systems risk embedding systematic bias into their predictions. In practice, this means that models trained on unreliable ground truth are prone to perpetuate measurement biases and for platform-agnostic AI vendors, encountering new types of endoscopes or lens settings may show that the AI models do not generalize well. Thus, the challenge is not only to train robust CNNs but also to ensure that the labels and imaging conditions underpinning them reflect true, reproducible size. Most validation studies, including the present work, try to mainly provide clinically relevant categorical binning (e.g. ≤5 mm, 6–9 mm, ≥10 mm) as the primary aim. This reflects clinical practice, where thresholds – particularly 10 mm – serve as critical decision points in surveillance guidelines. However, future AI models should ultimately provide continuous, millimeter accurate measurements rather than categorical binning. This task demands even greater rigor in ground truth data acquisition and AI development efforts. Training datasets will need sufficiently large numbers of polyps distributed across the full size spectrum, with enrichment around clinically decisive thresholds (e.g. 10 mm), to ensure they ultimately perform as intended in the real-world setting. Calibrated methods are indispensable for advancing the field. Laser-based endoscopic sizing systems represent one objective solution, though their adoption is limited by cost and hardware constraints [ 2 ] [ 4 ] [ 5 ] [ 6 ]. Post-resection measurement using calipers or on-site microscopy of fresh specimens provide another viable pathway, mitigating shrinkage artifacts and enabling reference sizing for AI training and validation. Such approaches are emerging as the most credible candidates for establishing reliable ground truth datasets. Importantly, the optimal ground truth method should allow for reproducibility checks and be confirmed by validation studies. Other recent studies, such as the work by Sudarevic et al. using waterjet-assisted AI systems, illustrate creative attempts to integrate adjunctive sizing markers directly into the AI measurement workflow [ 9 ]. However, such approaches still depend on non-calibrated adjunct instruments, while the ultimate goal remains AI models capable of inferring polyp size without requiring assistance from calibration instruments. Morphological cues such as pit pattern, vascular structures, and surface texture may provide the necessary information for models to achieve this without relying on additional tools in the field of view as shown in the current publication. Yet for AI systems to achieve broad clinical adoption, they must be benchmarked against ground truth derived from calibrated, verifiable standards. Antonelli et al. should be commended for taking the field forward by developing an AI-based sizing system that does not rely on additional tools in the field of view into clinical testing. Their work demonstrates feasibility, underscores the practical challenges of implementation, and highlights the value of real-world validation. At the same time, it points to the need for consensus across the field on how ground truth for polyp sizing is defined, measured, and validated. Without a reliable ground truth, the promise of AI-based sizing will be undermined by uncertainties in the very labels it seeks to approximate. Built and validated on reliable ground truth, however, AI will deliver the needed integrative solutions that go beyond detection and histology prediction, to include accurate AI-based sizing for colorectal polyps. Publication History Article published online: 02 October 2025 © 2025. Thieme. All rights reserved. Georg Thieme Verlag KG Oswald-Hesse-Straße 50, 70469 Stuttgart, Germany Refers to: Clinical implications of computer-aided real-time size estimation of colorectal polyps during colonoscopy: a prospective study Endoscopy eFirst DOI: 10.1055/a-2695-1978

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,018
score de la tête « metaresearch » (Gemma)0,072
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Théorique ou conceptuel · Signal consensuel: aucune
GenreSignal candidat: Commentaire · Signal consensuel: aucune
Score de désaccord entre enseignants0,018
Score d'incertitude au seuil0,096

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0180,072
Méta-épidémiologie (sens strict)0,0010,001
Méta-épidémiologie (sens large)0,0020,001
Bibliométrie0,0030,002
Études des sciences et des technologies0,0010,003
Communication savante0,0070,004
Science ouverte0,0030,002
Intégrité de la recherche0,0020,004
Charge utile insuffisante (le modèle a refusé de juger)0,0040,001

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,246
Tête enseignante GPT0,467
Écart entre enseignants0,221 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeThéorique ou conceptuel
Domainenon disponible
GenreCommentaire

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2025
Routes d'admission1
Résumé présentnon

Explorer davantage

Même revueEndoscopyMême sujetColorectal Cancer Screening and DetectionTravaux en français237 207