Accurate measurement of colonic polyps: In the “AI” of the beholder?
Bibliographic record
Abstract
Accurate sizing of polyps seen on colonoscopy is crucial to estimating colon cancer risk [1–3]. As such, international guidelines define surveillance intervals and recommended resection techniques based on the polyp size [4,5]. Unfortunately, estimation of lesion size in gastrointestinal endoscopy is fraught with human and technical inaccuracies that may lead to inappropriate surveillance recommendations in up to 30% of cases [6]. Artificial intelligence, in particular deep learning, which allows computers to ‘learn’ patterns without manual training has allowed computers to replace many traditionally human cognitive tasks in an objective and reproducible way [7]. Estimation of polyp size may be an ideal application of artificial intelligence given the historical challenge of this task documented in the gastroenterology literature. Several factors contribute to difficulties estimating the polyp size. The wide angle lens of the colonoscope creates a nonuniform magnification with relative compression in the periphery and disproportionate magnification in the centre of the field of view. Well described physician and patient factors also affect size estimates. Gender of the endoscopist, flatness and size of the polyp, as well as advanced patient age have all been correlated to the overestimation of polyp size and, interestingly, endoscopist experience does not seem to improve accuracy in size estimation across multiple studies [8,9]. Purpose-built endoscopic measurement devices have been developed and use of biopsy forceps as a measuring tool is commonly described [10]. However, dedicated measurement tools are not available in most endoscopy suites and can be onerous or expensive to use with every case. Due to magnification and image distortion, use of measurement devices can also be misleading if the device is not laid directly over the polyp or differs significantly from the polyp in size. This is often the case with biopsy forceps which have been shown to be significantly worse than visual estimation alone in one study [10]. Furthermore, measurement of the pathologic specimen as gold standard introduces small but significant error as well. Ignoring the obvious issue of piecemeal resection and coagulation artefact, en-bloc polyps may shrink when placed in formalin by as much as 15% compared to prefixation measurement [10]. Polypectomy alone can also cause in the range of 10% lesion shrinkage as confirmed by immediate prepolypectomy and postpolypectomy measurements taken with calibrated measurement tools prior to fixation [11]. Practically, endoscopists most often use their accumulated experience to ‘eyeball’ or make a visual estimate of the size of a lesion. While this has been an adequate approach to date, as expectations for quality progress, a more consistent approach to lesion sizing is welcome. Photograph editing programs that compensate for wide-angle lens distortion and magnification have been described in many articles [12]. Using these programs, an estimate of polyp size can be made based on the proportion of the visual field that the polyp occupies. Photographic measurement, however, requires a manual outline of a polyp which is time consuming and is not commonly used in practice. Similarly, compensation for magnification in polyp photographs requires knowledge of distance from the camera to the lesion to define the polyp size. Distance is typically measured using forceps or other working channel device of known length to contact the polyp, providing a minimal benefit over direct measurement with a dedicated device. Contactless methods of measurement have also been described using laser light to define size but require specialized equipment with the modest benefit [13]. To date, no specialized measurement technique has been widely adopted into everyday practice. Given its broad applications and dependence on real-time interpretation of video, gastrointestinal endoscopy has been the focus of rapidly expanding research into artificial intelligence applications [14]. One application of artificial intelligence, computer vision, has shown significant promise in its ability to enhance polyp detection and classification [15]. Computer vision is the ability of a computer to identify objects within an image, define them and assign importance to them in real time, which is to ‘see’ and ‘understand’ an image. Convolutional neural networks are a specific type of deep learning algorithm that is ideally suited to image interpretation. CNNs can be trained to recognize a class of objects such as polyps and outline or ‘segment’ them within the image. In this issue of the European Journal of Gastroenterology and Hepatology, Su et al. have made a valuable progress in estimating the lesion size consistently and efficiently by using artificial intelligence. In their article, they describe a simple though a rigorous mathematical model to account for endoscopic magnification and combine it with image segmentation via a convolutional neural network to first define and then measure colon polyps. In essence, they have trained a deep learning algorithm to outline a polyp and through a mathematical correction to ‘eyeball’ its size. Several unresolved issues and new questions arise from this preliminary work, however. For example, the authors fail to convince the reader of how robust their system would be in real life endoscopic practice. In particular, they fail to describe how to control for shooting distance of the endoscopic image, which was not described though presumably performed by either visual estimation or use of a measurement device in their study. Applications of the technique described by Su et al. beyond colonic polyp measurement would also be highly valuable and deserves to be explored, particularly in the measurement of critical endoscopic tasks such as stenting or Barrett’s surveillance. Importantly, few other published accounts of artificial intelligence-assisted polyp measurement exist. One commercially marketed endoscopy electronic health record system incorporates computer-assisted detection of polyps with an automated size estimate that requires contacting the polyp with a closed snare tip [16]. Unfortunately, the company does not list any publicly available research studies of the system on its website and no studies using the software were found on a manual search of medical literature databases and conference proceedings. An alternative approach to measurement was presented in abstract form in 2018 by Requa et al. [17] and describes a convolutional neural network that was able to reliably classify polyps in real time during colonoscopy into <6, 6–9 or >9 mm based on a training set of 7186 images using expert endoscopist labelling as a gold standard. Requa et al. [17] reported accuracies in their validation set of over 97%, an impressive number allowing for apparently consistent and objective measurement across colonoscopists if used, though by definition no more accurate than visual estimation by an expert endoscopist. An important consideration for researchers interested in computer-assisted lesion measurement is that while accurate sizing of polyps is crucial to human-guided risk assessment, it matters much less for adequately trained deep-learning algorithms. The reason being is that deep learning algorithms take into account size, pit pattern and countless other variables that the algorithm identifies as predictive without the requirement for manual input. The quintessence of a deep-learning algorithm is to extract what features are meaningful without manual programming and this may include dozens or hundreds of variables that a human operator would never take into account. However, while deep-learning may obviate the need for measurement in risk assessment of polyps, this in no way diminishes the value of accurate and efficient measurement. Accurate measurement is critical for technical planning and therapeutic decision making that artificial intelligence is unlikely to replace and polyp sizing is important for a global assessment of a patients’ colon cancer risk. Fast and accurate computer-assisted measurement of polyps should lead to better adherence to recommended surveillance guidelines and, eventually, to more meaningful sizing cutoffs in those guidelines themselves. If applied more broadly, efficient and accurate measurement also promises less waste and fewer errors in the seemingly banal, though critically important task of choosing the correct therapeutic consumable device. Fast, consistent automation of subjective tasks such as polyp measurement is an ideal application of artificial intelligence. While much work remains to be done, this is an important area which will no doubt be expanded upon. Acknowledgements Conflicts of interest M.F.B.is a CEO Satisfai Health, Founder ai4gi joint venture. Ai4gi has a co-development agreement with Olympus in artificial intelligence and colon polyps. There are no conflicts of interest for the remaining author.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.053 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.004 |
| Scholarly communication | 0.004 | 0.009 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.005 | 0.005 |
| Insufficient payload (model declined to judge) | 0.005 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".