Implementation of Artificial Intelligence-Assisted Endoscopy Across Canada—The CAG Artificial Intelligence Special Interest Group
Bibliographic record
Abstract
Artificial intelligence (AI) solutions use deep neural networks that imitate human brain neural interconnections. By doing so AI solutions can analyze endoscopy videos or images during or after an endoscopy (1). This allows feedback information to be superimposed on the endoscopy screen, assisting endoscopists, for example, during colonoscopy with the detection and differentiation of polyps, and prediction of histology or disease activity such as in inflammatory bowel disease. There are many other possible endoscopic applications. AI is expected to create many disruptive technologies in endoscopy with far-reaching and significant impact on quality and practice in general. Since endoscopy relies heavily on pattern recognition within video or image frames, AI seems an ideal technology to assist endoscopists in reaching quality benchmarks by providing live feedback. Colonoscopy quality metrics such as adenoma detection rate (ADR), withdrawal time of the colonoscope, cecal intubation rate and bowel cleanliness score are all ideal targets for automatic assessment and documentation using adapted AI solutions (1,2). Since these quality metrics are strongly linked to favorable patient outcomes, a positive clinical impact attributable to AI is to be expected. The first wave of AI applications that has been implemented in endoscopy units includes automated and integrated AI solutions to assist endoscopists in the detection and classification of colorectal polyps. Many image-based modalities have been developed to increase ADR, since an increase in ADR is closely linked to the prevention of interval colorectal cancers (CRC) (3). Although most image-based technologies have shown no or only moderate improvement in ADR, AI-based solutions have been shown to increase ADR by more than 10% (4) in robust international randomized controlled trials (RCTs). Table 1 summarizes the more than 10 available RCTs demonstrating an increase in ADR. This illustrates the tremendous potential of AI to impact the most important quality metric of endoscopic CRC prevention. Furthermore, the increase in ADR is independent of patient selection—that is, the benefits in increased ADR exist whether implemented for patients at average risk screening or following a positive fecal immunochemical test. More importantly, the increase in ADR is observed when using AI for both trainees as well as experienced endoscopists (5). Since an increase in ADR is closely linked to a reduction in interval cancer risk, the relevance of these findings cannot be overstated, and rapid integration of AI-based systems into clinical practice seems important. Meanwhile, almost all endoscopy platform providers offer integrated solutions for polyp detection, and platform-independent solutions are also now available, some having been approved by Health Canada. Results of adenoma detection from randomized controlled trials using AI ADR, adenoma detection rate; AI, artificial intelligence; CADe, computer aided polyp detection; HDWL, high defnition white light; ITT, intention to treat; OR, odds ratio; PPA, per protocol analysis; RR, relative risk; WL, white light. Results of adenoma detection from randomized controlled trials using AI ADR, adenoma detection rate; AI, artificial intelligence; CADe, computer aided polyp detection; HDWL, high defnition white light; ITT, intention to treat; OR, odds ratio; PPA, per protocol analysis; RR, relative risk; WL, white light. AI can characterize polyps and predict histology with high accuracy. While experienced endoscopists are usually able to predict pathology without needing AI, AI solutions can increase diagnostic accuracy for less experienced endoscopists (6). Studies have shown that when using AI, all endoscopists can meet critical accuracy benchmarks independent of their optical diagnosis skills, and perhaps sufficient for adopting a resect and discard colon polyp strategy. This will pave the way to replacing pathology for diminutive polyps with optical diagnosis, making colonoscopy practice more cost-effective (7). However, distinction between serrated or high-grade dysplastic pathology from other neoplastic polyps is, at present, not possible using AI. Future AI solutions will need to be trained to recognize this important granularity of polyp pathology subtypes (8). Once this is possible, widespread implementation of resect and discard and/or diagnose and leave strategies can be expected. Other AI solutions improve quality monitoring with automated assessment of quality metrics (such as bowel preparation, and completeness of the exam). For upper endoscopy, AI can assist in the assessment of the extent of Barrett’s esophagus, and in identifying areas with high-grade dysplasia or cancer (9). To improve productivity, workflow and quality control in endoscopy units, AI can create automated reports or automated analyses of databases, such as institutional or individual ADR (10). Development, research and regulatory approval have all allowed rapid progress in the development and dissemination of AI solutions. To further develop and promote AI technology, we have formed a special interest group (SIG) in AI at the Canadian Association of Gastroenterology (CAG. This CAG AI SIG core group is currently comprised of six gastroenterologists (the authors of this opinion piced) from five Canadian institutions across three provinces. We have started evaluating AI technologies using cohort studies and randomized controlled trials, and are in the process of establishing video and data biobanks to accrue raw data from which additional novel AI solutions can be created. Our research activities to date have been supported through seed funding from CAG, and we have organized and hosted webinars and sessions at Canadian Digestive Diseases Week (CDDW), inviting international experts on selected pertinent topics. Further activities of group members include the development and implementation of AI curricula since the next generation of gastroenterologists needs to be trained to develop and implement AI solutions at institutions across Canada. The CAG AI SIG has an open model inviting new members, industry and AI researchers to maximize the potential that this novel technology offers in improving endoscopy quality and patient outcomes. There was no funding received for this manuscript. D.R. has received research funding from ERBE Elektromedizin GmbH, Ventage, Pendopharm, Fuji and Pentax, and has received consultant or speaker fees from Boston Scientific Inc., ERBE Elektromedizin GmbH, and Pendopharm. M.F.B. is the CEO and Founder of Satisfai Health. A.B. is consultant for A.I. Vali. and Medctronic. All author authors have no conflicts of interest to declare.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.024 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.003 | 0.002 |
| Scholarly communication | 0.007 | 0.002 |
| Open science | 0.003 | 0.001 |
| Research integrity | 0.009 | 0.010 |
| Insufficient payload (model declined to judge) | 0.009 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".