Interobserver variation in tumor delineation of liver metastases using Magnetic Resonance Imaging
Bibliographic record
Abstract
<h2>Abstract</h2><h3>Background and purpose</h3> Magnetic Resonance Imaging (MRI) guided stereotactic body radiotherapy (SBRT) of liver metastases is an upcoming high-precision non-invasive treatment. Interobserver variation (IOV) in tumor delineation, however, remains a relevant uncertainty for planning target volume (PTV) margins. The aims of this study were to quantify IOV in MRI-based delineation of the gross tumor volume (GTV) of liver metastases and to detect patient-specific factors influencing IOV. <h3>Materials and methods</h3> A total of 22 patients with liver metastases from three primary tumor origins were selected (colorectal(8), breast(6), lung(8)). Delineation guidelines and planning MRI-scans were provided to eight radiation oncologists who delineated all GTVs. All delineations were centrally peer reviewed to identify outliers not meeting the guidelines. Analyses were performed both in- and excluding outliers. IOV was quantified as the standard deviation (SD) of the perpendicular distance of each observer's delineation towards the median delineation. The correlation of IOV with shape regularity, tumor origin and volume was determined. <h3>Results</h3> Including all delineations, average IOV was 1.6 mm (range 0.6–3.3 mm). From 160 delineations, in total fourteen single delineations were marked as outliers after peer review. After excluding outliers, the average IOV was 1.3 mm (range 0.6–2.3 mm). There was no significant correlation between IOV and tumor origin or volume. However, there was a significant correlation between IOV and regularity (Spearman's ρ<sub>s</sub> = -0.66; p = 0.002). <h3>Conclusion</h3> MRI-based IOV in tumor delineation of liver metastases was 1.3–1.6 mm, from which PTV margins for IOV can be calculated. Tumor regularity and IOV were significantly correlated, potentially allowing for patient-specific margin calculation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".