Deep Learning-based Detection and Tracking of Capsule Robots using Ultrasound Feedback in Gastrointestinal Tract
Bibliographic record
Abstract
Ingestible robotic capsules with locomotion capabilities and on‑board sampling mechanism have great potential for non‑invasive diagnostic and interventional use in the gastrointestinal tract. Real‑time tracking of capsule location and operational state is necessary for clinical application yet remains a significant challenge. To this end, this thesis investigates ultrasound-based detection and 3D tracking of capsule robots in the gastrointestinal tract by leveraging deep learning-based methods. As a pivotal step, an attention-based hierarchical deep learning approach is introduced. This method demonstrates the ability to simultaneously determine the mechanism state (i.e., the capsule is closed, open or lost) and in‑plane 2D pose of millimeter a capsule robot in ex-vivo tissue environments using ultrasound imaging. To train the neural networks, a representative dataset of the robotic capsule within ex‑vivo porcine stomachs is generated. The trained models achieve an accuracy of 96.9% for state classification and the mean estimation errors of 3.5 degrees and 1.76 mm for orientation and centroid position in ex-vivo experiments in porcine stomachs. However, conventional ultrasound B-mode imaging suffers from limited field of view, inability to image objects out of the scanning plane, and issues with low device visibility in echogenic in-vivo GI tract environments filled with bowel gas. To address these limitations and facilitate 3D tracking of capsule robots in tissue environments with intraluminal gas, an automatic robotic ultrasound tracking system for long-distance 3D tracking of a capsule robot in the GI tract is developed, with the ability to actively search for the lost capsule due to out-of-plane or out of the field of view motions. A hybrid deep learning model is proposed by combining transformer and convolutional neural networks for detecting and localizing the capsule robot. The attention mechanism built into the transformer enables the efficient capture of long-range capsule motions within the US image. The proposed system demonstrates continuous capsule tracking over 90 cm with a mean state detection accuracy of 90% and centroid localization accuracy of 1.5 mm across varying imaging parameters and artefact patterns.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".