Which failures do patient‐specific quality assurance systems need to catch?
Bibliographic record
Abstract
BACKGROUND: The Joint AAPM-ESTRO TG-360 is developing a quantitative framework to evaluate treatment verification systems used for patient-specific quality assurance (PSQA). A subgroup was commissioned to determine which potential failure modes had the greatest risk to treatment quality and safety, and therefore should be evaluated as part of the PSQA verification. PURPOSE: To create an extensive database of potential radiotherapy failure modes that should be detected by PSQA and to determine their relative importance for maximizing treatment quality. METHODS: The subgroup consisted of eight physicists from seven countries, including representatives from three international quality assurance groups. We collected error reports from RO-ILS, SAFRON, AAPM TG publications, and other literature, including international audits. We focused on the subset of failure modes that impact whether the planned dose matches the dose received by the patient. We performed a failure-mode-and-effects analysis (FMEA), estimating the severity (S), occurrence (O), and detectability (D) of each failure mode. Detectability was scored assuming that PSQA was not done but other routine clinical QA was performed, which allowed us to see the importance of PSQA for detecting each specific failure mode. We analyzed the risk priority number (RPN = O*S*D), O*S, and severity rankings to determine the priority of each failure mode. RESULTS: We collected 394 error reports, which we categorized into 33 failure modes that underwent FMEA. Five failure modes were in the top ranks for both RPN and O*S analysis: four involving treatment planning system (TPS) commissioning and one regarding patient model errors. The highest-ranking RPN failure modes were: TPS algorithm limitations, TPS commissioning errors [multileaf collimator (MLC) modeling, output factor, percent-depth-dose/tissue-maximum-ratio (PDD/TMR), off-axis factor], and patient weight variation. The highest O*S failure modes were similar, with the addition of external patient position variation and incorrect linear accelerator isocenter and cGy/monitor units calibration. RPN and O*S analyses prioritized failure modes that impacted multiple patients with high occurrence and detectability scores, while severity analysis gave higher priority to single-patient modes with high severity scores. The highest-ranking severity modes were MLC sequence deletion, collision, and TPS isocenter incorrect. CONCLUSION: We have developed a list of failure modes critical to be detected during PSQA and ranked them in order of importance. The top failure modes emphasize the importance of utilizing a variety of treatment verification systems for PSQA, from secondary dose calculation through in-vivo dosimetry, in order to detect all possible errors. For failure modes in the top quartile, PSQA is critical. Without adequate PSQA, these errors may go undetected unless caught by an external audit. This analysis can be useful for optimizing PSQA workflows and for designing evaluations of treatment verification systems, and will be used by the Joint AAPM-ESTRO TG-360 to determine an appropriate validation strategy.
Stored with the screening record, where it is evidence for the labels above.
How this classification was reachedexpand
The three-model screen
all 5,600 screened works →All three models called this out of scope.
Failure-mode analysis for radiotherapy patient-specific quality assurance; 'quality assurance' here is clinical dosimetry, not research practice (polysemy).
This study evaluates clinical radiotherapy quality assurance rather than research quality or practice.
Clinical radiotherapy patient-specific QA failure modes; quality-assurance polysemy, not research integrity or methods research.
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.099 | 0.263 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.007 | 0.006 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.005 | 0.006 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".