AAPM WGPE report 394: Simulated error training for the physics plan and chart review
Bibliographic record
Abstract
BACKGROUND: Simulated error training is a method to practice error detection in situations where the occurrence of error is low. Such is the case for the physics plan and chart review where a physicist may check several plans before encountering a significant problem. By simulating potentially hazardous errors, physicists can become familiar with how they manifest and learn from mistakes made during a simulated plan review. PURPOSE: The purpose of this project was to develop a series of training datasets that allows medical physicists and trainees to practice plan and chart reviews in a way that is familiar and accessible, and to provide exposure to the various failure modes (FMs) encountered in clinical scenarios. METHODS: A series of training datasets have been developed that include a variety of embedded errors based on the risk-assessment performed by American Association of Physicists in Medicine (AAPM) Task Group 275 for the physics plan and chart review. The training datasets comprise documentation, screen shots, and digital content derived from common treatment planning and radiation oncology information systems and are available via the Cloud-based platform ProKnow. RESULTS: Overall, 20 datasets have been created incorporating various software systems (Mosaiq, ARIA, Eclipse, RayStation, Pinnacle) and delivery techniques. A total of 110 errors representing 50 different FMs were embedded with the 20 datasets. The project was piloted at the 2021 AAPM Annual Meeting in a workshop where participants had the opportunity to review cases and answer survey questions related to errors they detected and their perception of the project's efficacy. In general, attendees detected higher-priority FMs at a higher rate, though no correlation was found between detection rate and the detectability of the FMs. Familiarity with a given system appeared to play a role in detecting errors, specifically when related to missing information at different locations within a given software system. Overall, 96% of respondents either agreed or strongly agreed that the ProKnow portal and training datasets were effective as a training tool, and 75% of respondents agreed or strongly agreed that they planned to use the tool at their local institution. CONCLUSIONS: The datasets and digital platform provide a standardized and accessible tool for training, performance assessment, and continuing education regarding the physics plan and chart review. Work is ongoing to expand the project to include more modalities, radiation oncology treatment planning and information systems, and FMs based on emerging techniques such as auto-contouring and auto-planning.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.044 | 0.109 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.013 | 0.008 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".