Artificial intelligence and colon capsule endoscopy: Automatic detection of ulcers and erosions using a convolutional neural network
Bibliographic record
Abstract
BACKGROUND AND AIM: Colon capsule endoscopy (CCE) has become a minimally invasive alternative for conventional colonoscopy. Nevertheless, each CCE exam produces between 50 000 and 100 000 frames, making its analysis time-consuming and prone to errors. Convolutional neural networks (CNNs) are a type of artificial intelligence (AI) architecture with high performance in image analysis. This study aims to develop a CNN model for the identification of colonic ulcers and erosions in CCE images. METHODS: A CNN model was designed using a database of CCE images. A total of 124 CCE exams performed between 2010 and 2020 in two centers were reviewed. For CNN development, a total of 37 319 images were extracted, 33 749 showing normal colonic mucosa and 3570 showing colonic ulcers and erosions. Datasets for CNN training, validation, and testing were created. The performance of the algorithm was evaluated regarding its sensitivity, specificity, positive and negative predictive values, accuracy, and area under the curve. RESULTS: The network had a sensitivity of 96.9% and a specificity of 99.9% specific for the detection of colonic ulcers and erosions. The algorithm had an overall accuracy of 99.6%. The area under the curve was 1.00. The CNN had an image processing capacity of 90 frames per second. CONCLUSIONS: The developed algorithm is the first CNN-based model to accurately detect ulcers and erosions in CCE images, also providing a good image processing performance. The development of these AI systems may contribute to improve both the diagnostic and time efficiency of CCE exams, facilitating CCE adoption to routine clinical practice.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".