Detecting lumbar lesions in <sup>99m</sup>Tc‐MDP SPECT by deep learning: Comparison with physicians
Bibliographic record
Abstract
Abstract Purpose 99mTc‐MDP single‐photon emission computed tomography (SPECT) is an established tool for diagnosing lumbar stress, a common cause of low back pain (LBP) in pediatric patients. However, detection of small stress lesions is complicated by the low quality of SPECT, leading to significant interreader variability. The study objectives were to develop an approach based on a deep convolutional neural network (CNN) for detecting lumbar lesions in 99mTc‐MDP scans and to compare its performance to that of physicians in a localization receiver operating characteristic (LROC) study. Methods Sixty‐five lesion‐absent (LA) 99mTc‐MDP studies performed in pediatric patients for evaluating LBP were retrospectively identified. Projections for an artificial focal lesion were acquired separately by imaging a 99mTc capillary tube at multiple distances from the collimator. An approach was developed to automatically insert lesions into LA scans to obtain realistic lesion‐present (LP) 99mTc‐MDP images while ensuring knowledge of the ground truth. A deep CNN was trained using 2.5D views extracted in LP and LA 99mTc‐MDP image sets. During testing, the CNN was applied in a sliding‐window fashion to compute a 3D “heatmap” reporting the probability of a lesion being present at each lumbar location. The algorithm was evaluated using cross‐validation on a 99mTc‐MDP test dataset which was also studied by five physicians in a LROC study. LP images in the test set were obtained by incorporating lesions at sites selected by a physician based on clinical likelihood of injury in this population. Results The deep learning (DL) system slightly outperformed human observers, achieving an area under the LROC curve (AUCLROC) of 0.830 (95% confidence interval [CI]: [0.758, 0.924]) compared with 0.785 (95% CI: [0.738, 0.830]) for physicians. The AUCLROC for the DL system was higher than that of two readers (difference in AUCLROC [ΔAUCLROC] = 0.049 and 0.053) who participated to the study and slightly lower than that of two other readers (ΔAUCLROC = −0.006 and −0.012). Another reader outperformed DL by a more substantial margin (ΔAUCLROC = −0.053). Conclusion The DL system provides comparable or superior performance than physicians in localizing small 99mTc‐MDP positive lumbar lesions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".