Comparative evaluation of Ordered Subset Expectation Maximization and Bayesian Penalized Likelihood algorithms for PET/CT image reconstruction in various malignancies using 18F-FDG and 68Ga-PSMA-11 tracers
Bibliographic record
Abstract
Purpose This study compares the Ordered Subset Expectation Maximization (OSEM) and Bayesian Penalized Likelihood (BPL) algorithms for Positron Emission Tomography/Computed Tomography (PET/CT) image reconstruction using 18F- fluorodeoxyglucose (FDG) and 68Ga-labeled Prostate-Specific Membrane Antigen(68Ga-PSMA-11) tracers. Methods A retrospective analysis was conducted on 33 patients with various malignancies, including 25 undergoing 18F-FDG PET/CT scans and 8 undergoing 68Ga-PSMA-11 PET/CT scans. Scans were reconstructed using both OSEM and BPL (Q.Clear) algorithms, evaluating key metrics such as Standardized Uptake Value (SUV)max, SUVpeak, background SUV, and tumor-to-background ratio (TBR). Results Thirty-three patients (mean age: 67.53 ± 11.78 years) with 100 lesions (80 FDG, 20 PSMA) were analyzed. In the FDG group, significant differences were observed in lesion SUVpeak, liver SUVpeak, SD of liver SUVmean, bladder SUVmean, SD of bladder SUVmean, and TBR, with BPL generally producing higher values except for liver SUVpeak, SD of liver SUVmean, and SD of bladder SUVmean. In the PSMA group, BPL enhanced most metrics except for the SD of liver SUVmean and the SD of bladder SUVmean. While strong linear correlations between BPL and OSEM metrics were observed (Pearson r>0.85 for most parameters), Bland-Altman analysis revealed wide limits of agreement, particularly for TBR in the PSMA group (-13.00 to 23.89), indicating substantial variability in absolute values between methods. Assessment of the relationship between lesion volume and SUVmax differences (BPL–OSEM) revealed a weak, non-significant negative correlation in the total cohort(r = – 0.14, p = 0.16) and in the FDG subgroup(r= – 0.14, p = 0.24). Conclusion Both reconstruction methods demonstrate clinical utility, with BPL producing statistically higher values for several quantitative metrics, such as SUVmax and TBR, without markedly improving lesion detectability. While strong correlations were observed between BPL and OSEM values, the wide limits of agreement, particularly for TBR in PSMA imaging, suggest these methods may not be directly interchangeable in longitudinal studies. Harmonization strategies may help reduce inter-method variability and improve scan comparability in longitudinal or multicenter settings. Prospective approaches, such as reconstruction-specific reference ranges or scaling factors, could further support harmonization efforts in clinical trials. For longitudinal monitoring, consistent use of the same reconstruction method is recommended to ensure reliable quantification.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.019 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".