Comment on: Diagnosis of giant cell arteritis: reply
Bibliographic record
Abstract
We thank Dr Ing for his interest and comments [1] on our review paper entitled ‘Diagnosis of giant cell arteritis’ (GCA) [2]. We will provide a structured rebuttal to his criticisms following the same order presented in his letter. Given the review nature of our article, we aimed to critically describe the most important aspects in the diagnosis of GCA, with a particular focus on the current advances in this field. TAB is already a very well-established diagnostic modality for GCA, offering many clinical advantages mentioned in our article (e.g. high diagnostic specificity, differential diagnosis with other diseases, potential prognostic value). However, to the best of our knowledge, the published diagnostic sensitivity for TAB in patients with GCA shows great disparity, ranging from 39 to 95% [3, 4], which means that it is correct to say its ‘sensitivity can be as low as 39%’, referencing the TABUL study [3]. Although we acknowledge this study had several limitations, not only reflected in the unsatisfactory performance characteristics of TAB, but also of ultrasound to diagnose GCA [5], it was the first international multicentre study comparing the clinical effectiveness and cost-effectiveness of both diagnostic modalities using a reference standard diagnosis for GCA. It included a high number of patients with suspected GCA (n = 430), recruited from 20 different sites, and was able to provide an improved diagnostic accuracy using a combined strategy with ultrasound or TAB, together with clinical judgement. The TABUL study was innovative, reflected the clinical reality of the time in which it was conducted (2010–2013), with all its inherent flaws, proved the importance of always taking into account the clinical manifestations of the disease in a diagnostic approach to patients with suspected GCA and thus should not be neglected when reviewing the evidence for diagnosing GCA. Nevertheless, we appreciate Dr Ing’s mention of the recent meta-analysis showing a pooled estimate sensitivity of 77% for TAB [6]. It is unquestionably an important article, also not without its natural shortcomings, but it had not been published before our review was submitted.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.074 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.003 | 0.004 |
| Scholarly communication | 0.003 | 0.007 |
| Open science | 0.004 | 0.003 |
| Research integrity | 0.049 | 0.039 |
| Insufficient payload (model declined to judge) | 0.006 | 0.008 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".