An International Study of Factors Affecting Variability of Dosimetry Calculations, Part 2: Overall Variabilities in Absorbed Dose
Bibliographic record
Abstract
Dosimetry for personalized radiopharmaceutical therapy has gained considerable attention. Many methods, tools, and workflows have been developed to estimate absorbed dose (AD). However, standardization is still required to reduce variability of AD estimates across centers. One effort for standardization is the Society of Nuclear Medicine and Molecular Imaging <sup>177</sup>Lu Dosimetry Challenge, which comprised 5 tasks (T1–T5) designed to assess dose estimate variability associated with the imaging protocol (T1 vs. T2 vs. T3), segmentation (T1 vs. T4), time integration (T4 vs. T5), and dose calculation (T5) steps of the dosimetry workflow. The aim of this work was to assess the overall variability in AD calculations for the different tasks. <b>Methods:</b> Anonymized datasets consisting of serial planar and quantitative SPECT/CT scans, organ and lesion contours, and time-integrated activity maps of 2 patients treated with <sup>177</sup>Lu-DOTATATE were made available globally for participants to perform dosimetry calculations and submit their results in standardized submission spreadsheets. The data were carefully curated for formal mistakes and methodologic errors. General descriptive statistics for ADs were calculated, and statistical analysis was performed to compare the results of different tasks. Variability in ADs was measured using the quartile coefficient of dispersion. <b>Results:</b> ADs to organs estimated from planar imaging protocols (T2) were lower by about 60% than those from pure SPECT/CT (T1), and the differences were statistically significant. Importantly, the average differences in dose estimates when at least 1 SPECT/CT acquisition was available (T1, T3, T4, T5) were within ±10%, and the differences with respect to T1 were not statistically significant for most organs and lesions. When serial SPECT/CT images were used, the quartile coefficients of dispersion of ADs for organs and lesions were on average less than 20% and 26%, respectively, for T1; 20% and 18%, respectively, for T4 (segmentations provided); and 10% and 5%, respectively, for T5 (segmentation and time-integrated activity images provided). <b>Conclusion:</b> Variability in ADs was reduced as segmentation and time-integration data were provided to participants. Our results suggest that SPECT/CT-based imaging protocols generate more consistent and less variable results than planar imaging methods. Effort at standardizing segmentation and fitting should be made, as this may substantially reduce variability in ADs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".