Fully automated coronary artery calcium score and risk categorization from chest CT using deep learning and multiorgan segmentation: A validation study from National Lung Screening Trial (NLST)
Bibliographic record
Abstract
Background: The National Lung Screening Trial (NLST) has shown that screening with low dose CT in high-risk population was associated with reduction in lung cancer mortality. These patients are also at high risk of coronary artery disease, and we used deep learning model to automatically detect, quantify and perform risk categorisation of coronary artery calcification score (CACS) from non-ECG gated Chest CT scans. Materials and methods: Automated calcium quantification was performed using a neural network based on Mask regions with convolutional neural networks (R-CNN) for multiorgan segmentation. Manual evaluation of calcium was carried out using proprietary software. This study used 80 patients to train the segmentation model and randomly selected 1442 patients were used for the validation of the algorithm. We compared the model generated results with Ground Truth. Results: Automatic cardiac and aortic segmentation model worked well (Mean Dice score: 0.91). Cohen's kappa coefficient between the reference actual and the interclass computed predictive categories on the test set is 0.72 (95 % CI: 0.61-0.83). Our method correctly classifies the risk group in 78.8 % of the cases and classifies the subjects in the same group. F-score is measured as 0.78; 0.71; 0.81; 0.82; 0.92 in calcium score categories 0(CS:0), I (1-99), II (100-400), III (400-1000), IV (>1000), respectively. 79 % of the predictive scores lie in the same categories, 20 % of the predictive scores are one category up or down, and only 1.2 % patients were more than one category off. For the presence/absence of coronary artery calcifications, our deep learning model achieved a sensitivity of 90 % and a specificity of 94 %. Conclusion: Fully automated model shows good correlation compared with reference standards. Automating the process could improve diagnostic ability, risk categorization, facilitate primary prevention intervention, improve morbidity and mortality, and decrease healthcare costs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".