Myocardial perfusion SPECT radiomic features reproducibility assessment: Impact of image reconstruction and harmonization
Bibliographic record
Abstract
BACKGROUND: Coronary artery disease (CAD) has one of the highest mortality rates in humans worldwide. Single-photon emission computed tomography (SPECT) myocardial perfusion imaging (MPI) provides clinicians with myocardial metabolic information non-invasively. However, there are some limitations to interpreting SPECT images performed by physicians or automatic quantitative approaches. Radiomics analyzes images objectively by extracting quantitative features and can potentially reveal biological characteristics that the human eye cannot detect. However, the reproducibility and repeatability of some radiomic features can be highly susceptible to segmentation and imaging conditions. PURPOSE: We aimed to assess the reproducibility of radiomic features extracted from uncorrected MPI-SPECT images reconstructed with 15 different settings before and after ComBat harmonization, along with evaluating the effectiveness of ComBat in realigning feature distributions. MATERIALS AND METHODS: A total of 200 patients (50% normal and 50% abnormal) including rest and stress (without attenuation and scatter corrections) MPI-SPECT images were included. Images were reconstructed using 15 combinations of filter cut-off frequencies, filter orders, filter types, reconstruction algorithms, number of iterations and subsets resulting in 6000 images. Image segmentation was performed on the left ventricle in the first reconstruction for each patient and applied to 14 others. A total of 93 radiomic features were extracted from the segmented area, and ComBat was used to harmonize them. The intraclass correlation coefficient (ICC) and overall concordance correlation coefficient (OCCC) tests were performed before and after ComBat to examine the impact of each parameter on feature robustness and to assess harmonization efficiency. The ANOVA and the Kruskal-Wallis tests were performed to evaluate the effectiveness of ComBat in correcting feature distributions. In addition, the Student's t-test, Wilcoxon rank-sum, and signed-rank tests were implemented to assess the significance level of the impacts made by each parameter of different batches and patient groups (normal vs. abnormal) on radiomic features. RESULTS: Before applying ComBat, the majority of features (ICC: 82, OCCC: 61) achieved high reproducibility (ICC/OCCC ≥ 0.900) under every batch except Reconstruction. The largest and smallest number of poor features (ICC/OCCC < 0.500) were obtained by IterationSubset and Order batches, respectively. The most reliable features were from the first-order (FO) and gray-level co-occurrence matrix (GLCM) families. Following harmonization, the minimum number of robust features increased (ICC: 84, OCCC: 78). Applying ComBat showed that Order and Reconstruction were the least and the most responsive batches, respectively. The most robust families, in a descending order, were found to be FO, neighborhood gray-tone difference matrix (NGTDM), GLCM, gray-level run length matrix (GLRLM), gray-level size zone matrix (GLSZM), and gray-level dependence matrix (GLDM) under Cut-off, Filter, and Order batches. The Wilcoxon rank-sum test showed that the number of robust features significantly differed under most batches in the Normal and Abnormal groups. CONCLUSION: The majority of radiomic features show high levels of robustness across different OSEM reconstruction parameters in uncorrected MPI-SPECT. ComBat is effective in realigning feature distributions and enhancing radiomic features reproducibility.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".