Artificial intelligence-powered coronary artery disease diagnosis from SPECT myocardial perfusion imaging: a comprehensive deep learning study
Bibliographic record
Abstract
BACKGROUND: Myocardial perfusion imaging (MPI) using single-photon emission computed tomography (SPECT) is a well-established modality for noninvasive diagnostic assessment of coronary artery disease (CAD). However, the time-consuming and experience-dependent visual interpretation of SPECT images remains a limitation in the clinic. PURPOSE: We aimed to develop advanced models to diagnose CAD using different supervised and semi-supervised deep learning (DL) algorithms and training strategies, including transfer learning and data augmentation, with SPECT-MPI and invasive coronary angiography (ICA) as standard of reference. MATERIALS AND METHODS: A total of 940 patients who underwent SPECT-MPI were enrolled (281 patients included ICA). Quantitative perfusion SPECT (QPS) was used to extract polar maps of rest and stress states. We defined two different tasks, including (1) Automated CAD diagnosis with expert reader (ER) assessment of SPECT-MPI as reference, and (2) CAD diagnosis from SPECT-MPI based on reference ICA reports. In task 2, we used 6 strategies for training DL models. We implemented 13 different DL models along with 4 input types with and without data augmentation (WAug and WoAug) to train, validate, and test the DL models (728 models). One hundred patients with ICA as standard of reference (the same patients in task 1) were used to evaluate models per vessel and per patient. Metrics, such as the area under the receiver operating characteristics curve (AUC), accuracy, sensitivity, specificity, precision, and balanced accuracy were reported. DeLong and pairwise Wilcoxon rank sum tests were respectively used to compare models and strategies after 1000 bootstraps on the test data for all models. We also compared the performance of our best DL model to ER's diagnosis. RESULTS: In task 1, DenseNet201 Late Fusion (AUC = 0.89) and ResNet152V2 Late Fusion (AUC = 0.83) models outperformed other models in per-vessel and per-patient analyses, respectively. In task 2, the best models for CAD prediction based on ICA were Strategy 3 (a combination of ER- and ICA-based diagnosis in train data), WoAug InceptionResNetV2 EarlyFusion (AUC = 0.71), and Strategy 5 (semi-supervised approach) WoAug ResNet152V2 EarlyFusion (AUC = 0.77) in per-vessel and per-patient analyses, respectively. Moreover, saliency maps showed that models could be helpful for focusing on relevant spots for decision making. CONCLUSION: Our study confirmed the potential of DL-based analysis of SPECT-MPI polar maps in CAD diagnosis. In the automation of ER-based diagnosis, models' performance was promising showing accuracy close to expert-level analysis. It demonstrated that using different strategies of data combination, such as including those with and without ICA, along with different training methods, like semi-supervised learning, can increase the performance of DL models. The proposed DL models could be coupled with computer-aided diagnosis systems and be used as an assistant to nuclear medicine physicians to improve their diagnosis and reporting, but only in the LAD territory. CLINICAL TRIAL NUMBER: Not applicable.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".