Provider Attitudes and Perceptions on Using Artificial Intelligence in Colonoscopy: A Systematic Review and Meta-Analysis
Bibliographic record
Abstract
Background and Aims: Colonoscopy is the gold standard screening modality for colorectal cancer; however, it is operator-dependent and reliant on exam quality. Incorporating artificial intelligence (AI) into colonoscopy may improve adenoma detection and clinical outcomes, but this is a sociotechnical challenge that requires effective human-AI teaming incorporating provider attitudes. Methods: We conducted a systematic review of studies evaluating attitudes and perspectives of providers toward AI-assisted colonoscopy. Participant responses to outcome questions of interest were combined across the studies to calculate pooled proportion (Pp) and 95% confidence interval (CI). Top-ranked perceived advantages and disadvantages in each study were defined as the items that >50% of the study participants voted for. Results: Out of 2044 abstracts screened, 13 studies were included representing 1538 providers who were mostly gastroenterologists or trainees and 25%-100% had direct experience using AI in a clinical setting. Overall, a large majority were interested in using AI (Pp = 80%, 95% CI 70%-89%, n = 8 studies) and believed it can improve adenoma or polyp detection rate (Pp = 74%, 95% CI 68%-80%, n = 4 studies). Among 5 studies addressing financial implications, about half were concerned about the cost of using AI (52%, 95% CI 24%-79%). An average of 38% of respondents (95% CI 9%-73%) from 4 studies raised concern regarding accountability for misdiagnosis. High number of false positives and an absence of clinical guidelines were top-ranked perceived disadvantages in 2 studies. Conclusion: Most gastroenterology providers expressed interest in using AI systems with colonoscopy and believed it can improve adenoma detection rate. Cost, high number of false positives, and lack of professional society guidelines were among top perceived concerns.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.005 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".