Evaluation of artificial intelligence-based treatment planning and delineation algorithms for external radiotherapy
Bibliographic record
Abstract
External radiotherapy aims to treat cancer cells using ionizing radiation. The challenge is to irradiate target volumes while sparing healthy organs. These treatments rely on delineating regions of interest and treatment planning, carried out through Treatment Planning Systems (TPS). These processes are manually realized, demand precision and time. Their quality and execution time depend on the operator. In this context, automation solutions emerge, promising time efficiency, practices uniformization, and maintaining or enhancing treatment quality. This thesis aims to evaluate and clinically implement 2 artificial intelligence (AI) algorithms, one for automatic segmentation and the other for treatment planning. These algorithms are available in RayStation TPS (RaySearch Medical Laboratories AB, Stockholm, Sweden). The first chapter introduces the clinical environment for implementing automation techniques. The second chapter concerns the automatic segmentation algorithm. It is quantitatively and qualitatively evaluated against reference structures validated by physician. Initially, the AI is compared to 2 other algorithms within RS: one based on multi-atlases, and another on statistical models. Subsequently, AI validation is performed for 24 different organs within the thorax, abdomen, head, and neck. Lastly, this AI is compared to a second AI available in another automatic segmentation software called Limbus (Limbus AI Inc., Regina, SK, Canada). The third chapter focuses on treatment planning automation using AI. First, we describe the construction of three databases (DBs) provided to RaySearch, enabling the creation of AI models tailored to our clinical practices for pelvic and whole-brain treatments. This AI is compared to manually generated plans and a multi-criteria optimization (MCO) algorithm, demonstrating AI's clinical utility. Models are adjusted using a cohort of 100 patients, validating acceptability and deliverability of automatically generated plans for pelvic locations. However, these plans are improvable and can be enhanced manually by operators. We present, later in this chapter, a user experience with clinical use of AI for the first 79 automatically generated and manually improved treatment plans. The chapter's end focuses on the value of training AI with a database generated by MCO plans. AI evaluation, compared to manually generated and MCO plans, reveals that AI-generated plans do not match MCO plan quality but enhance manual plans. In the final chapter, we propose a guide for implementing automatic methods clinically. This ranges from defining expectations to ethical AI use, utilizing on-site developed Python scripts, development, validation methods, and limitations encountered for each automation type. In summary, the algorithms assessed in this thesis, notably the AIs, save time, standardize practices, and aid treatment planning quality enhancement. However, their implementation demands substantial preliminary effort, and their utilization should be combined with human expertise to ensure treatment quality.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".