Computer aided detection and diagnosis of polyps in adult patients undergoing colonoscopy: a living clinical practice guideline
Bibliographic record
Abstract
Abstract Clinical question In adult patients undergoing colonoscopy for any indication (screening, surveillance, follow-up of positive faecal immunochemical testing, or gastrointestinal symptoms such as blood in the stools) what are the benefits and harms of computer-aided detection (CADe)? Context and current practice Colorectal cancer (CRC), the third most common cancer and the second leading cause of cancer-related death globally, typically arises from adenomatous polyps. Detection and removal of polyps during colonoscopy can reduce the risk of cancer. CADe systems use artificial intelligence (AI) to assist endoscopists by analysing real-time colonoscopy images to detect potential polyps. Despite their increasing use in clinical practice, guideline recommendations that carefully balance all patient-important outcomes remain unavailable. In this first iteration of a living guideline, we address the use of CADe at the level of an individual patient. Evidence Evidence for this recommendation is drawn from a living systematic review of 44 randomised controlled trials (RCTs) involving more than 30 000 participants and a companion microsimulation study simulating 10 year follow-up for 100 000 individuals aged 60-69 years to assess the impact of CADe on patient-important outcomes. While no direct evidence was found for critical outcomes of colorectal cancer incidence and post-colonoscopy cancer incidence, low certainty data from the trials indicate that CADe may increase positive endoscopy findings. The microsimulation modelling, however, suggests little to no effect on CRC incidence, CRC-related mortality, or colonoscopy-related complications (perforation and bleeding) over the 10 year follow-up period, although low certainty evidence indicates CADe may increase the number of colonoscopies performed per patient. A review of values and preferences identified that patients value mortality reduction and quality of care but worry about increased anxiety, overdiagnosis, and more frequent surveillance. Recommendation For adults who have agreed to undergo colonoscopy, we suggest against the routine use of CADe (weak recommendation). How this guideline was created An international panel, including three patient partners, 11 healthcare providers, and seven methodologists, deemed by MAGIC and The BMJ to have no relevant competing interests, developed this recommendation. For this guideline the panel took an individual patient approach. The panel started by defining the clinical question in PICO format, and prioritised outcomes including CRC incidence and mortality. Based on the linked systematic review and microsimulation study, the panel sought to balance the benefits, harms, and burdens of CADe and assumed patient preferences when making this recommendation Understanding the recommendation The guideline panel found the benefits of CADe on critical outcomes, such as CRC incidence and post-colonoscopy cancer incidence, over a 10 year follow up period to be highly uncertain. Low certainty evidence suggests little to no impact on CRC-related mortality, while the potential burdens—including more frequent surveillance colonoscopies—are likely to affect many patients. Given the small and uncertain benefits and the likelihood of burdens, the panel issued a weak recommendation against routine CADe use. The panel acknowledges the anticipated variability in values and preferences among patients and clinicians when considering these uncertain benefits and potential burdens. In healthcare settings where CADe is available, individual decision making may be appropriate. Updates This is the first iteration of a living practice guideline. The panel will update this living guideline if ongoing evidence surveillance identifies new CADe trial data that substantially alters our conclusions about CRC incidence, mortality, or burdens, or studies that increase our certainty in values and preferences of individual patients. Updates will provide recommendations on the use of CADe from a healthcare systems perspective (including resource use, acceptability, feasibility, and equity), as well as the combined use of CADe and computer aided diagnosis (CADx). Users can access the latest guideline version and supporting evidence on MAGICapp, with updates periodically published in The BMJ .
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".