Sensitivity and Specificity of Radiographic Scoring Instruments for Detecting Change in Axial Psoriatic Arthritis
Bibliographic record
Abstract
OBJECTIVE: There is no widely recognized method used to assess axial disease in psoriatic arthritis (PsA). We aimed to determine the sensitivity to change of the Bath Ankylosing Spondylitis Radiology Index for the spine (BASRI-s), the modified Stoke Ankylosing Spondylitis Spine Score (mSASSS), the Radiographic Ankylosing Spondylitis Spine Score (RASSS), and the PsA Spondylitis Radiology Index (PASRI) in axial PsA. METHODS: Radiographs of 105 patients with axial PsA were retrieved for 2 time points at least 2 years apart and subsequently anonymized. All radiographs were scored by 3 rheumatologists blinded to name and order of examination using an electronic application that allowed recording of disease manifestations specific to axial PsA and automatically calculated the BASRI-s, mSASSS, RASSS, and PASRI scores. An independent expert determined whether there was true radiographic progression from an overall impression after viewing the radiographs with knowledge of chronologic order. The sensitivity, specificity, and odds ratios for every 1-unit increase in the scores were determined to identify true change. RESULTS: Of the patients studied, 25 (24%) showed progression, as determined by the independent expert. The respective sensitivity and specificity values for an increase in score to detect true change were as follows: 0.48 and 0.78 (BASRI-s), 0.52 and 0.84 (mSASSS), 0.44 and 0.84 (RASSS), and 0.52 and 0.74 (PASRI). Logistic regression analyses showed that an increase of 1 point in the respective scores was associated with the following odds ratios for identifying true progression: BASRI-s 3.0, mSASSS 5.27, RASSS 3.70, and PASRI 3.06. CONCLUSION: Available scoring systems for quantifying radiographic axial PsA have moderate sensitivity but high specificity for detecting true change.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.024 | 0.086 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".