Computer-Aided Diagnosis for the Resect and Discard Strategy for Colorectal Polyps: A Systematic Review and Meta-Analysis
Bibliographic record
Abstract
Aims According to the Resect and Discard strategy, endoscopists can replace post-polypectomy pathology with real-time prediction (optical diagnosis) of polyp histology during colonoscopy. This strategy is only applicable to small polyps≤5mm and if the endoscopist prediction was made with high confidence. The variability in real-time optical diagnosis among different endoscopists can be standardized by the high accuracy expected from computer-aided diagnosis systems (CADx). The aim of this meta-analysis is to provide a preliminary estimate of the accuracy of CADx and to assess its effect in clinical practice. Methods We conducted a search of MEDLINE, EMBASE, and Scopus databases, covering studies published from inception to October 31, 2023. We included histologically-verified accuracy diagnostic studies that evaluated the real-time optical diagnosis performance of endoscopists for polyps≤5mm in the entire colon. The study had two main endpoints: 1) to assess the accuracy of CADx-alone, including sensitivity, specificity, and negative (NPV) and positive (PPV) predictive values; and 2) to compare CADx-unassisted and -assisted diagnosis in terms of proportion of polyps resected and discarded and and the appropriateness of the surveillance intervals according to ASGE/ESGE guidelines. For this second endpoint, diagnoses were restricted to only high-confidence diagnosis as required by the guidelines. Results We analyzed 8 studies using 5 different CADx systems (1,850 patients with 3,815 polyps≤5mm). The CADx-alone pooled sensitivity and NPV were 87.6% (95% CI: 0.813 – 0.920) and 85% (95% CI: 0.751 – 0.915), respectively. While specificity and PPV were 83.8% (95% CI: 0.746 – 0.901) and 86.5% (95% CI: 0.825 – 0.897), respectively. When limiting our analysis to high-confidence diagnosis, the proportion of diminutive polyps that would have been resected and discarded appears to be 90% (95% CI: 0.81 – 0.95) and 91% (95% CI: 0.89 – 0.93) in the unassisted and assisted arms, respectively. In detail, sensitivity and specificity were, respectively, 93.2% (95% CI: 0.901 – 0.953) and 79.4% (95% CI: 0.538 – 0.928) for the unassisted arm and 91.9% (95% CI: 0.852 – 0.957) and 87.7% (95% CI: 0.754 – 0.943) for the assisted arm. Conclusions The CADx-alone strategy demonstrated very good optical diagnosis performance; however, the analysis identified certain outliers. This underscores the importance of thoroughly investigating a CADx system before its utilization. A limitation of our study is that not all CADx systems used were regulatory approved, and it's noteworthy that most participating endoscopists had extensive experience in optical diagnosis. Publication History Article published online: 15 April 2024 © 2024. European Society of Gastrointestinal Endoscopy. All rights reserved. Georg Thieme Verlag KG Rüdigerstraße 14, 70469 Stuttgart, Germany
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.019 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.009 | 0.015 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".