Retrospective Comparative Analysis of Diagnostic Tools for Giant Cell Arteritis
Bibliographic record
Abstract
Objectives To assess the diagnostic performance of various giant cell arteritis (GCA) diagnostic tools including temporal artery biopsy (TAB), which is the current gold standard for GCA diagnosis, temporal artery ultrasound (US), presence of concurrent polymyalgia rheumatica (PMR) and the American College of Rheumatology (ACR) criteria for GCA diagnosis. Methods A retrospective chart review was conducted on 138 GCA cases between 2006-2024 where 1 or more GCA diagnostic tests were conducted. True positives (TP), true negatives (TN), false positives (FP), and false negatives (FN) were determined using the final clinical diagnoses and were used to calculate sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV) and accuracy. Relative risks (RR) of PMR history and PMR symptoms at presentation for GCA diagnosis were also calculated. Results TAB showed an accuracy of 0.722, a sensitivity of 0.449 and a specificity of 1 (n=97). There was no significant difference in diagnostic accuracy of TAB when the biopsy specimen was greater than 1cm compared to when it was less than 1cm (p=0.262). US showed an accuracy of 0.834, a sensitivity of 0.636 and a specificity of 0.95 (n=31). RR of GCA diagnosis in patients with a history of PMR was 1.36 and in patients presenting with both GCA and PMR symptoms was 1.48 (n=136). ACR criteria showed an accuracy of 0.553, a sensitivity of 0.744 and a specificity of 0.609 (n=85). Conclusion The results of our study support that TAB may have great specificity but low sensitivity in diagnosis of GCA. We also found that US may have similar diagnostic performance to TAB in diagnosing GCA, indicating its utility as an effective non-invasive diagnostic tool. Our findings also indicate a lower sensitivity and specificity for the ACR criteria compared to previous studies. Additionally, the relative risk analysis underscores the increased likelihood of GCA diagnosis in patients with a history of PMR or presenting with concurrent PMR symptoms. Given the variability introduced by smaller sample sizes, this study should be considered a preliminary investigation into the clinical utility of these diagnostic tools. Future research with larger sample sizes is essential to validate these results and to better define the interplay of TAB, US, ACR criteria and PMR history in GCA diagnosis. This will help determine the optimal diagnostic approach and whether alternative tests might effectively supplement or replace TAB in clinical practice.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.013 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".