International Harmonization of Technical Approaches to Kidney SABR – An International Radiosurgery Consortium of the Kidney (IROCK) Contouring Project
Bibliographic record
Abstract
Stereotactic ablative body radiotherapy (SABR) is an emerging treatment for patients with primary renal cell carcinoma (RCC), however variation in treatment protocols can exist between institutions. The goals of this study were to measure the variation in contouring RCC tumors for patients being treated with SABR and to develop consensus recommendations. An international panel of 16 radiation oncologists was created from the IROCK meeting during ASTRO 2023. Four patient cases were: Case 1, a renal tumor greater than 10 cm in size with an IVC tumor thrombus; Case 2, a central renal tumor abutting the renal hilum; Case 3, a local recurrence of RCC post-nephrectomy; and Case 4, a residual tumor post-radiofrequency ablation (RFA). For each Case, panelists were asked for radiation planning details and to contour the target volumes on representative axial images using a computer-based training tool. Comparison of panelist contours with the were performed using the Dice-Similarity Coefficient (DSC), the Mean Distance to Agreement (MDA) and the Hausdorff Distance (HD). The DSC measures the overlap between two contours, so a higher DSC suggests greater agreement. The MDA and HD represent the mean and maximum distances between points on the two contours, so higher MDA and HD represent lower agreement. Consensus target volumes were derived using the STAPLE algorithm and discussed amongst the panel. Altogether, the panel included radiation oncologists from Canada, the USA, Australia, the Netherlands, and India. All panelists had previously treated at least 10 patients with SABR for primary RCC. Table 1 shows the DSC, MDA and HD for each case. Using an ANOVA analysis, for all Cases, the DSC, MDA and HD were not statistically different between participants (p = 0.32, p = 0.24, and p = 0.23, respectively). On qualitative inspection of participant contours, Case 4 showed the most agreement, followed by Case 3, Case 2, and then Case 1. There was good agreement from our international expert panel on contouring renal tumors in all four Cases, with Case 4 having the greatest agreement, and Case 1 having the least agreement. Consensus recommendations based on this study may help improve the quality of SABR for RCC moving forward.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.081 | 0.029 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.004 | 0.001 |
| Open science | 0.004 | 0.005 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".