MétaCan
Menu
Back to cohort
Record W4403828500 · doi:10.1002/mp.17468

Which failures do patient‐specific quality assurance systems need to catch?

2024· article· en· W4403828500 on OpenAlexaff
J OˈDaniel, Víctor Hernández, Catharine H. Clark, M. Esposito, Jöerg Lehmann, Andrea McNiven, I. Olaciregui-Ruiz, Stephen F. Kry

Bibliographic record

VenueMedical Physics · 2024
Typearticle
Languageen
FieldPhysics and Astronomy
TopicAdvanced Radiotherapy Techniques
Canadian institutionsBaker Hughes (Canada)University of Toronto
Fundersnot available
KeywordsQuality assuranceFailure mode and effects analysisMedicineMedical physicsAuditQuality (philosophy)Reliability engineeringComputer scienceAccountingEngineeringBusiness

Abstract

BACKGROUND: The Joint AAPM-ESTRO TG-360 is developing a quantitative framework to evaluate treatment verification systems used for patient-specific quality assurance (PSQA). A subgroup was commissioned to determine which potential failure modes had the greatest risk to treatment quality and safety, and therefore should be evaluated as part of the PSQA verification. PURPOSE: To create an extensive database of potential radiotherapy failure modes that should be detected by PSQA and to determine their relative importance for maximizing treatment quality. METHODS: The subgroup consisted of eight physicists from seven countries, including representatives from three international quality assurance groups. We collected error reports from RO-ILS, SAFRON, AAPM TG publications, and other literature, including international audits. We focused on the subset of failure modes that impact whether the planned dose matches the dose received by the patient. We performed a failure-mode-and-effects analysis (FMEA), estimating the severity (S), occurrence (O), and detectability (D) of each failure mode. Detectability was scored assuming that PSQA was not done but other routine clinical QA was performed, which allowed us to see the importance of PSQA for detecting each specific failure mode. We analyzed the risk priority number (RPN = O*S*D), O*S, and severity rankings to determine the priority of each failure mode. RESULTS: We collected 394 error reports, which we categorized into 33 failure modes that underwent FMEA. Five failure modes were in the top ranks for both RPN and O*S analysis: four involving treatment planning system (TPS) commissioning and one regarding patient model errors. The highest-ranking RPN failure modes were: TPS algorithm limitations, TPS commissioning errors [multileaf collimator (MLC) modeling, output factor, percent-depth-dose/tissue-maximum-ratio (PDD/TMR), off-axis factor], and patient weight variation. The highest O*S failure modes were similar, with the addition of external patient position variation and incorrect linear accelerator isocenter and cGy/monitor units calibration. RPN and O*S analyses prioritized failure modes that impacted multiple patients with high occurrence and detectability scores, while severity analysis gave higher priority to single-patient modes with high severity scores. The highest-ranking severity modes were MLC sequence deletion, collision, and TPS isocenter incorrect. CONCLUSION: We have developed a list of failure modes critical to be detected during PSQA and ranked them in order of importance. The top failure modes emphasize the importance of utilizing a variety of treatment verification systems for PSQA, from secondary dose calculation through in-vivo dosimetry, in order to detect all possible errors. For failure modes in the top quartile, PSQA is critical. Without adequate PSQA, these errors may go undetected unless caught by an external audit. This analysis can be useful for optimizing PSQA workflows and for designing evaluations of treatment verification systems, and will be used by the Joint AAPM-ESTRO TG-360 to determine an appropriate validation strategy.

Stored with the screening record, where it is evidence for the labels above.

How this classification was reachedexpand

The three-model screen

all 5,600 screened works →

All three models called this out of scope.

stratum: aff_core · design weight: 5595.24 (the sample is stratified; any rate computed without the weight is wrong)
Claude Opus 4.8OUT
genre: empirical
about Canada: no
confidence: high

Failure-mode analysis for radiotherapy patient-specific quality assurance; 'quality assurance' here is clinical dosimetry, not research practice (polysemy).

GPT-5.6 (high)OUT
genre: empirical
about Canada: no
confidence: high

This study evaluates clinical radiotherapy quality assurance rather than research quality or practice.

Grok 4.5OUT
genre: empirical
about Canada: no
confidence: high

Clinical radiotherapy patient-specific QA failure modes; quality-assurance polysemy, not research integrity or methods research.

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.099
metaresearch head score (Gemma)0.263
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.099
Threshold uncertainty score0.000

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0990.263
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0070.006
Science and technology studies0.0020.002
Scholarly communication0.0050.006
Open science0.0020.003
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.013
GPT teacher head0.305
Teacher spread0.292 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designTheoretical or conceptual
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations15
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueMedical PhysicsSame topicAdvanced Radiotherapy TechniquesFrench-language works237,207