Low dose positron emission tomography emulation from decimated high statistics: A clinical validation study
Bibliographic record
Abstract
PURPOSE: The fundamental nature of positron emission tomography (PET), as an event detection system, provides some flexibility for data handling, including retrospective data manipulation. The reorganization of acquisition data allows the emulation of new scans arising from identical radiotracer spatial distributions, but with different statistical compositions, and is especially useful for evaluating the stability and reproducibility of reconstruction algorithms or when investigating extremely low count conditions. This approach is ubiquitous in the research literature but has only been validated, from the point of view of the noise properties, with numerical simulations and phantom data. We present here the first experiment comparing PET images of the same human subjects generated with two separate injections of radiotracer, using actual low dose (LD) data to validate a randomly decimated emulation from a standard dose scan. A key point of the work is focused on the randoms fractions, which scale differently than the trues at varying activity levels. METHODS: Eleven patients with non-small cell lung cancer were enrolled in the study. Each imaging session consisted of two independent FDG-PET/CT scans: a LD scan followed by a standard dose (SD) scan. Images were first reconstructed, using filtered back-projection (FBP) and OSEM incorporating time-of-flight information and point-spread function modeling (PSFTOF), from the LD and SD datasets comprising all counts from each scanned bed position. The number of true counts was recorded for all LD scans, and independent, count-matched emulations (ELD) were reconstructed from the SD data. Noise distribution within the liver and standardized uptake value reproducibility within a population of contoured, tracer-avid lesion volumes were evaluated across scans and statistics. RESULTS: The randoms fraction estimates were 17.4 ± 1.6% (14.9-19.4) in the LD data and 42 ± 2.3% (37.1-45.5) in the SD data. Eleven lesions were identified and volumes of interest were generated with a 50% threshold isocontour for each lesion, in every image. The distributions of metabolic volumes, means and maxima defined by the contoured volumes-of-interest (VOIs) were similar between the LD and SD sets. A two-tailed, matched t-test was performed on the populations of region statistics for both LD and ELD reconstructions, and the t-statistics were 1.1 (P = 0.267) and -0.22 (P = 0.828) for the background liver VOIs and -0.54 (P = 0.603) and 0.23 (P = 0.821) for the lesion VOIs, for FBP and PSFTOF respectively. In every test, the null hypothesis that the two populations had the same mean could not be rejected at the 5% significance level. CONCLUSIONS: Our results demonstrate that clinical LD PET scans can indeed be accurately emulated by the statistical decimation of standard dose scans, and this was achieved through validation by images generated with unbiased random coincidence estimations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".