Enhancing diagnosis of gout with deep learning in dual-energy computed tomography: a retrospective analysis of crystal and artefact differentiation
Bibliographic record
Abstract
OBJECTIVES: To evaluate whether the application of deep learning (DL) could achieve high diagnostic accuracy in differentiating between green colour coding, indicative of tophi, and clumpy artefacts observed in dual-energy computed tomography (DECT) scans. METHODS: A comprehensive analysis of 18,704 regions of interest (ROIs) extracted from green foci in DECT scans obtained from 47 patients with gout and 27 gout-free controls was performed. The ROIs were categorized into three size groups: small, medium, and large. Convolutional neural network (CNN) analysis on a per-lesion basis and support vector machine (SVM) analysis on a per-patient basis were performed. The area under the receiver operating characteristic curve, sensitivity, specificity, positive predictive value, and negative predictive value of the models were compared. RESULTS: For small ROIs, the sensitivity and specificity of the CNN model were 81.5% and 96.1%, respectively; for medium ROIs, 82.7% and 96.1%, respectively; for large ROIs, 91.8% and 86.9%, respectively. Additionally, the DL algorithm exhibited accuracies of 88.5%, 88.6%, and 91.0% for small, medium, and large ROIs, respectively. In the per-patient analysis, the SVM approach demonstrated a sensitivity of 87.2%, a specificity of 100%, and an accuracy of 91.8% in distinguishing between patients with gout and gout-free controls. CONCLUSION: Our study demonstrates the effectiveness of the DL algorithm in differentiating between green colour coding indicative of crystal deposition and clumpy artefacts in DECT scans. With high sensitivity, specificity, and accuracy, the utilization of DL in DECT for diagnosing gout enables precise lesion classification, facilitating early-stage diagnosis and promoting timely intervention approaches.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".