Multicenter quantitative 18F-fluorodeoxyglucose positron emission tomography performance harmonization: use of hottest voxels towards more robust quantification
Bibliographic record
Abstract
Background: Harmonization methods reduce variability between different make and models of positron emission tomography (PET) scanners. The study aims to explore harmonization strategies that lead to comparable and robust quantitative metrics in a multicenter setting. Methods: NEMA IEC Phantom data acquisition was performed for low and high spheres-to-background ratios (SBR4:1 and 10:1) on six PET/CT (computed tomography) scanners. Different reconstruction sets, including the number of sub-iterations, number of subsets, and full width at half maximum (FWHM) for each scanner, were evaluated towards optimized and harmonized reconstruction settings. Recovery coefficients (RCs) of four quantitative metrics, including standardized uptake value (SUV)max, SUVISO-50 (SUVmean in 50% isocontour), SUVpeak, and mean uptake of 10 highest concentration voxels were evaluated as RCmax, RCISO-50, RCpeak, and RC10V, representing percent difference relative to the static ground truth case as functions of sphere sizes. A set of image reconstruction parameters was proposed for harmonized reconstruction to minimize variability between scanners. The root mean square error (RMSE), curvature, and reproducibility were examined. The proposed reconstruction protocols for harmonization and standard clinical reconstruction settings were compared to each other across all scanners. Results: A significant difference (P value <0.0001) was observed in the aforementioned quantitative metrics between SBR10 and SBR4. Reconstruction parameter sets with the smallest RMSE and RC values within 10% bias were identified as the best candidate for harmonization. The coefficient of variation of the mean value of RCs (CVMRC) shows a remarkable reduction of about 28%, 26%, 32%, and 19% in harmonized reconstruction settings for MRCmax, MRCISO-50, MRCpeak, and MRC10V, respectively. CVMRC for MRC10V in the harmonized reconstruction setting was 5.9% in SBR4, while the smallest value in SBR10 belongs to MRCpeak, with a value of 5.8%. The reproducibility of RC is improved by deriving the value from ten hottest voxels and is equally reproducible with RCpeak. Compared to RCmax and RCISO-50, the variability is reduced by 18% and 22% if ten voxels are pooled. Conclusions: Harmonizing PET/CT systems with and without point spread function/time of flight (PSF/TOF) using various vendor-developed image reconstruction algorithms improves the quantification reproducibility. RC10V, likewise RCpeak, is superior to the rest of the quantitative indices in terms of accuracy and reproducibility and helpful in quantifying lesion volume below 1 mL.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".