Submillimeter accuracy in radiosurgery is not possible
Bibliographic record
Abstract
Arguing against the Proposition is Sonja Dieterich, Ph.D. After completing her Ph.D. in Nuclear Physics at Rutgers University in 2002, Dr. Dieterich received training in Medical Physics at Georgetown University Hospital, Washington, DC, from 2002 to 2003. In 2003, she accepted a faculty position at Georgetown. From 2007 to 2012, she worked at Stanford University Hospital as Clinical Associate Professor and Chief of Radiosurgery Physics. Since April 2012 she is an Associate Professor and Physics Residency Co-Director at the University of California Davis. Dr. Dieterich is Chair of the AAPM Task Group 135 "QA for Robotic Radiosurgery" and a member of the ASTRO Physics and Multi-Disciplinary QA Committees. Her current research interests are the development of QA/QM programs for new technologies, image-guided brachytherapy, and Veterinary Radiation Oncology. The advent of cranial radiosurgical therapy has allowed a nonsurgical approach to the treatment of cranial lesions.1 The obvious benefits apply to both malignant as well as benign lesions. Historically a rigid invasive frame, attached to the patient's skull via small screws, acted as both immobilizer and localizer. With such devices, the frame defines a 3D coordinate system within which the skull and targeted lesion are intended to remain fixed from the time of the planning CT scan to the completion of treatment. More modern radiosurgical systems may use 3D "image-guided" positioning for exact alignment. Accuracy of radiosurgery has often been considered as ultimately relying on the rigidity of the immobilizer, as opposed to the treatment system as a whole. This ignores multiple other factors that must also be included. Typical quoted values for frame-based immobilization accuracy are 0.3–0.6 mm.2–4 This measure relates only to flex between the frame and the skull, and the stability of the treatment apparatus. The system accuracy as a whole must also include end-to-end uncertainties that result from the spatial resolution of the original CT, resolution and linearity of the MRI scan, contouring uncertainty, treatment planning grid resolution, algorithm errors, CT to MRI registration uncertainty, etc. The now-dated AAPM report 54 (Ref. 5) suggested that the total uncertainty can reach 2–3 mm. The planning system is designed to calculate an optimized dose distribution around a physician-determined region of interest (ROI). Since the physician-drawn ROI is a function of the window and level differences on CT and MRI, there is variability between physicians, which even for well-identified normal structures (let alone often less distinct targets) can exceed 1–2 mm.6 A recent study by a Canadian group demonstrated a shift in tumor isocenter of 1.4 mm by simply altering the time of CT imaging after injected contrast.7 An analysis of the registration accuracy of CT with MRI for several algorithms determined typical errors along the x, y, and z axes of approximately 0.68, 1.04, and 0.60 mm, respectively. When added in quadrature this amounts to a vector of 1.4 mm.8 Others have published even greater variability.9 A recent Gamma Knife end-to-end study, evaluating the distance to agreement, determined that the uncertainty was better than expected based on quadrature-sum but still over 1 mm.10 The planning system dose grid is unlikely to have a resolution of less than 1 mm and the dose calculation algorithm will itself contain some uncertainty in placement of dose within the grid due to imperfect modeling of radiation transport. Since a significant percent of target definition is done on MRI datasets it is important to note that even for MRI systems compliant with ACR guidelines, the distortion can reach up to 2 mm.11 Taking all of the above uncertainties into account is essential in deriving a realistic overall accuracy for SRS treatment. It is apparent that radiosurgery treatment using current technology and considering human factors, when observed from end-to-end, cannot reach submillimeter accuracy. The crux of this debate is the definition of "Accuracy in Radiosurgery." In the literature, there is significant confusion caused by imprecise language. If we are debating the mechanical accuracy of the delivery system to align to isocenter, the answer is straightforward and supported by literature. Current mechanical engineering techniques meet this standard easily.3,12,13 Frameless, image-guided radiosurgery adds the requirement to match the imaging isocenter to the mechanical isocenter. A look at the QA tables contained in TG-142 confirms that the expert authors consider this an achievable goal for standard QA.14 The next level of achieving submillimeter accuracy is to match the radiation isocenter to the mechanical isocenter using image-guidance. This task can be achieved by, for example, using an end-to-end (E2E) test. A literature search confirms submillimeter accuracy in E2E tests is achievable on CyberKnife,3 Truebeam,15 and Gamma Knife13 machines. We should assume similar E2E accuracy for other delivery systems (to be published soon). Radiosurgery targets are not always spherical or elliptical. Patient-specific delivery QA (DQA) is done to assess how mechanical, imaging, treatment planning, and radiation isocenter uncertainties combine in a rigid phantom. In IMRT, gamma criteria of 3%/3 mm are customarily used. To argue my case, I must argue that a reduction to 3%/1 mm is possible.15 This is the point where the argument becomes too complex to support a binary answer to our hypothesis. First, the 1 mm criterion. Everyone who has performed a DQA is aware that working on an accuracy scale of 1 mm is a very time-intensive undertaking. The two QA devices we have available to reach this resolution are film and gels. Other QA devices have been shown to be unable to detect mechanical shifts of 1 mm in delivery.16 Therefore, we cannot use them to verify that a radiosurgery device is accurate to <1 mm. What about the 3% dose uncertainty: does it belong in this argument? I could choose the easy way out and argue that we are mixing units, hence it should be excluded. But as I like to point out, any uncertainty in dose, may it originate from dose calculation algorithms, delivery uncertainties, or any other possible cause, is equivalent to moving an isodose line. Up to this point, we have been working in the physics realm of phantom testing for which we have confirmation of our ability to achieve submillimeter accuracy, so this proves the Proposition to be invalid. Once we step out of the phantom world and apply radiosurgery to patients, I have to concede. Even in as seemingly simple a task as contouring, we are still limited by imaging, causing contouring uncertainties of several millimeters among expert physicians. Not to mention residual patient motion or changes in internal anatomy. Since this is a physics publication, I am going to claim these factors are beyond this discussion, and still claim to have disproved the Proposition. What an interesting approach to this debate. My esteemed colleague initially claims that 1 mm accuracy is achievable, then executes an elegant demi-tour and concedes that it cannot be met, only to follow with another demi-tour, discounting any parameters that are inconvenient, and claims victory! I applaud Dr. Dieterich's breakdown of the components of radiosurgery. My colleague states that, when these parameters are applied to a phantom, using certain specialized tools, assuming that a 3% dose error is nonexistent, ignoring contouring uncertainties, ignoring organ motion, ignoring imaging uncertainties, etc. we can achieve 1 mm accuracy! The reasoning is that this is a medical physics publication; hence clinical realities that affect accuracy can be excluded from the equation. Not so I say! After all, as medical physicists we can hardly modify the accuracy of our linear accelerator's isocenter. This is determined by a vendor over which we have little control. The accuracy of our imaging system, our treatment planning system, registration software, etc. are also predetermined for us. We can measure their accuracy, but cannot modify them. Similarly, we can measure the clinical uncertainties that exist, from one clinician to another in contouring, for example, or how windowing levels affect contouring, or organ motion during respiration. We can measure those parameters but not affect them. They are still part of the equation. This reality of accuracy should hold whether the debate takes place in Medical Physics or in the International Journal of Radiation Oncology, Biology, Physics. Because physicists concentrate largely on parameters that are within their control and can be readily applied to a phantom, such as physical or dosimetric measurements, we sometimes forget, or choose to ignore, the big picture. Our physician colleagues hear us tell them that our treatments are accurate within a certain tolerance, and they believe us. We need to step back and look at the reality from a larger perspective that includes all relevant parameters. The weakest links, be they imaging, mechanical, dosimetric, or clinical, must be identified and addressed. When we do take these parameters into consideration, submillimeter accuracy in a patient is currently not achievable. Dr. Bichay states: "It is apparent that radiosurgery treatment using current technology and considering human factors, when observed from end-to-end, cannot reach sub-millimeter accuracy." This is in direct contradiction to peer-reviewed papers3,15,17–19 demonstrating submillimeter accuracy with end-to-end tests for several radiosurgery methods. Furthermore, he is offering no supporting literature for his statement that the planning system dose grids are unlikely to have dose grid resolutions of less than 1 mm. Modern treatment planning systems have dose grid resolutions matching (Accuray MultiPlan) or exceeding20 the voxel resolution of the CT image. At 512 × 512 with a field of view of 250 mm for a cranial scan, this is equivalent to 0.5 mm. The interpolation between dose points pushes the resolution even higher, because for photon treatments dose gradients are relatively smooth. It is also no longer true that a significant percentage of target definition is done on MR data sets alone. The fraction of radiosurgery patients with plans generated on MR-based planning systems has been declining; more widely used modalities such as linac based SRS use CT based treatment planning, with MR being used as an additional, complementary component of the contouring process. Even where MR is used for target definition (but not for localization), the 2 mm MR distortion Dr. Bichay references is measured over the extent of the MR phantom (>100 mm). Over the typical diameter of a large brain metastasis, 20– 30 mm, the expected distortion is less than 0.5 mm. I do agree with Dr. Bichay that there is significant variability in defining the target volume. Quantitative imaging of disease has not reached the quality or resolution needed to achieve submillimeter accuracy. While I disagree with Dr. Bichay when I maintain that we can achieve submillimeter accuracy spatially and most likely dosimetrically (at least for rigid targets in homogeneous areas of the body), I do agree with him that submillimeter accuracy in contouring is as yet beyond our reach.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.018 | 0.038 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.013 |
| Scholarly communication | 0.004 | 0.010 |
| Open science | 0.003 | 0.006 |
| Research integrity | 0.006 | 0.011 |
| Insufficient payload (model declined to judge) | 0.006 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".