Learning-based autonomous navigation, benchmark environments and simulation framework for endovascular interventions
Bibliographic record
Abstract
Endovascular interventions are a life-saving treatment for many diseases, but they suffer from drawbacks such as radiation exposure and the potential scarcity of proficient physicians. Robotic assistance during these interventions could be a promising support for these problems. Research focusing on autonomous endovascular interventions using artificial intelligence-based methodologies is gaining popularity. However, variability in assessment environments hinders the comparability of different approaches, primarily due to each study employing a unique evaluation framework. In this study, we present autonomous endovascular instrument navigation based on deep reinforcement learning for three distinct digital benchmark interventions: BasicWireNav, ArchVariety, and DualDeviceNav. The benchmarks focus on aortic arch to supra-aortic navigation, representing fundamental large-vessel navigation skills. The benchmark interventions were implemented with our modular simulation framework stEVE (simulated EndoVascular Environment). Autonomous controllers were trained solely in simulation and evaluated in simulation and on physical test benches with camera and fluoroscopy feedback. Autonomous control for BasicWireNav and ArchVariety reached success rates up to 98/100 in simulation and was successfully transferred to the physical test benches with a success rate of up to 97/100. The experiments demonstrate the feasibility of stEVE and its potential to transfer simulation-trained controllers to real-world scenarios. However, they also reveal areas that offer opportunities for future research. Furthermore, this work reduces barriers to entry and increases the comparability of research on learning-based assistance systems for endovascular navigation by providing open-source training scripts, benchmarks, and the stEVE framework. • Novel benchmark environments for autonomous endovascular navigation using stEVE. • Successful simulation-to-reality transfer with 98% to 97% success rate. • Multi-instrument coordination challenges identified in DualDeviceNav benchmark. • Open-source framework enables reproducible endovascular robotics research. • Reward design ablation shows two-component combinations accelerate learning.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".