Method of computing direction-dependent margins for the development of consensus contouring guidelines
Bibliographic record
Abstract
BACKGROUND: Clinical target volume (CTV) contouring guidelines are frequently developed through studies in which experts contour the CTV for a representative set of cases for a given treatment site and the consensus CTVs are analyzed to generate margin recommendations. Measures of interobserver variability are used to quantify agreement between experts. In cases where an isotropic margin is not appropriate, however, there is no standard method to compute margins in specified directions that represent possible routes of tumor spread. Moreover, interobserver variability metrics are often measures of volume overlap that do not account for the dependence of disagreement on direction. To aid in the development of consensus contouring guidelines, this study demonstrates a novel method of quantifying CTV margins and interobserver variability in clinician-specified directions. METHODS: The proposed algorithm was applied to 11 cases of non-spine bone metastases to compute the consensus CTV margin in each direction of intraosseous and extraosseous disease. The median over all cases for each route of spread yielded the recommended margins. The disagreement between experts on the CTV margin was quantified by computing the median of the coefficients of variation for intraosseous and extraosseous margins. RESULTS: The recommended intraosseous and extraosseous margins were 7.0 mm and 8.0 mm, respectively. The median coefficient of variation quantifying the margin disagreement between experts was 0.59 and 0.48 for intraosseous and extraosseous disease. CONCLUSIONS: The proposed algorithm permits the generation of margin recommendations in relation to adjacent anatomy and quantifies interobserver variability in specified directions. This method can be applied to future consensus CTV contouring studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".