Which failures do patient‐specific quality assurance systems need to catch?
Notice bibliographique
Résumé
BACKGROUND: The Joint AAPM-ESTRO TG-360 is developing a quantitative framework to evaluate treatment verification systems used for patient-specific quality assurance (PSQA). A subgroup was commissioned to determine which potential failure modes had the greatest risk to treatment quality and safety, and therefore should be evaluated as part of the PSQA verification. PURPOSE: To create an extensive database of potential radiotherapy failure modes that should be detected by PSQA and to determine their relative importance for maximizing treatment quality. METHODS: The subgroup consisted of eight physicists from seven countries, including representatives from three international quality assurance groups. We collected error reports from RO-ILS, SAFRON, AAPM TG publications, and other literature, including international audits. We focused on the subset of failure modes that impact whether the planned dose matches the dose received by the patient. We performed a failure-mode-and-effects analysis (FMEA), estimating the severity (S), occurrence (O), and detectability (D) of each failure mode. Detectability was scored assuming that PSQA was not done but other routine clinical QA was performed, which allowed us to see the importance of PSQA for detecting each specific failure mode. We analyzed the risk priority number (RPN = O*S*D), O*S, and severity rankings to determine the priority of each failure mode. RESULTS: We collected 394 error reports, which we categorized into 33 failure modes that underwent FMEA. Five failure modes were in the top ranks for both RPN and O*S analysis: four involving treatment planning system (TPS) commissioning and one regarding patient model errors. The highest-ranking RPN failure modes were: TPS algorithm limitations, TPS commissioning errors [multileaf collimator (MLC) modeling, output factor, percent-depth-dose/tissue-maximum-ratio (PDD/TMR), off-axis factor], and patient weight variation. The highest O*S failure modes were similar, with the addition of external patient position variation and incorrect linear accelerator isocenter and cGy/monitor units calibration. RPN and O*S analyses prioritized failure modes that impacted multiple patients with high occurrence and detectability scores, while severity analysis gave higher priority to single-patient modes with high severity scores. The highest-ranking severity modes were MLC sequence deletion, collision, and TPS isocenter incorrect. CONCLUSION: We have developed a list of failure modes critical to be detected during PSQA and ranked them in order of importance. The top failure modes emphasize the importance of utilizing a variety of treatment verification systems for PSQA, from secondary dose calculation through in-vivo dosimetry, in order to detect all possible errors. For failure modes in the top quartile, PSQA is critical. Without adequate PSQA, these errors may go undetected unless caught by an external audit. This analysis can be useful for optimizing PSQA workflows and for designing evaluations of treatment verification systems, and will be used by the Joint AAPM-ESTRO TG-360 to determine an appropriate validation strategy.
Conservé avec la notice de tri, où il sert de preuve aux étiquettes ci-dessus.
Comment cette classification a été obtenuedéplier
Le tri à trois modèles
les 5 600 travaux triés →Les trois modèles l'ont jugé hors champ.
Failure-mode analysis for radiotherapy patient-specific quality assurance; 'quality assurance' here is clinical dosimetry, not research practice (polysemy).
This study evaluates clinical radiotherapy quality assurance rather than research quality or practice.
Clinical radiotherapy patient-specific QA failure modes; quality-assurance polysemy, not research integrity or methods research.
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,099 | 0,263 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,002 |
| Bibliométrie | 0,007 | 0,006 |
| Études des sciences et des technologies | 0,002 | 0,002 |
| Communication savante | 0,005 | 0,006 |
| Science ouverte | 0,002 | 0,003 |
| Intégrité de la recherche | 0,001 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,002 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».