Validation of the Southend giant cell arteritis probability score in a Scottish single-centre fast-track pathway
Bibliographic record
Abstract
OBJECTIVE: The aim was to provide external validation of the Southend GCA probability score (GCAPS) in patients attending a GCA fast-track pathway (GCA FTP) in NHS Lanarkshire. METHODS: Consecutive GCA FTP patients between November 2018 and December 2020 underwent GCAPS assessment as part of routine care. GCA diagnoses were supported by US of the cranial and axillary arteries (USS), with or without temporal artery biopsy (TAB), and confirmed at 6 months. Percentages of patients with GCA according to GCAPS risk group, performance of total GCAPS in distinguishing GCA/non-GCA final diagnoses, and test characteristics using different GCAPS binary cut-offs were assessed. Associations between individual GCAPS components and GCA and the value of USS and TAB in the diagnostic process were also explored. RESULTS: Forty-four of 129 patients were diagnosed with GCA, including 0 of 41 GCAPS low-risk patients (GCAPS <9), 3 of 40 medium-risk patients (GCAPS 9-12) and 41 of 48 high-risk patients (GCAPS >12). Overall performance of GCAPS in distinguishing GCA/non-GCA was excellent [area under the receiver operating characteristic curve, 0.976 (95% CI 0.954, 0.999)]. GCAPS cut-off ≥10 had 100.0% sensitivity and 67.1% specificity for GCA. GCAPS cut-off ≥13 had the highest accuracy (91.5%), with 93.2% sensitivity and 90.6% specificity. Several individual GCAPS components were associated with GCA. Sensitivity of USS increased by ascending GCAPS risk group (nil, 33.3% and 90.2%, respectively). TAB was diagnostically useful in cases where USS was inconclusive. CONCLUSION: This is the first published study to describe application of GCAPS outside the specialist centre where it was developed. Performance of GCAPS as a risk stratification tool was excellent. GCAPS might have additional value for screening GCA FTP referrals and guiding empirical glucocorticoid treatment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".