Bibliographic record
Abstract
Abstract End-to-End (E2E) testing is a method originating from computer science that is designed to determine whether an application communicates as required with hardware, networks, databases, and other applications. This paper is to advocate that the quality management (QM) of modern radiation therapy (RT) would benefit from more regular use of E2E based quality assurance (QA) in the local clinic. The argument is that modern RT delivery is performed through some process linked by a chain of interdependent stages and actions mediated by complex interchanges during the patient’s treatment. These actions along the chain are often modified due to decisions by clinical staff who are interpreting information acquired along the process. While physics QA can validate that each of these steps are technically achievable (e.g., through machine QA) such conventional QA does not guarantee that the overall process is being carried out as planned even when it has been described by a well-defined protocol and delivered by well-trained staff. The paper briefly reviews the changes in programmatic design as RT has become more complex, the associated changes in RT QM, and some past examples of E2E testing in RT clinics, usually performed during the implementation of some new RT technique or during external audits of the clinic’s practice. The paper then makes the case for increased E2E QA based on the lessons learned from this experience and ends with some suggestions for implementing effective and sustainable E2E testing in a clinic’s QM program.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".