Metric Properties of the SPARCC Score of the Sacroiliac Joints — Data from Baseline, 3-month, and 12-month Followup in the SPACE Cohort
Bibliographic record
Abstract
OBJECTIVE: To evaluate metric properties of the SpondyloArthritis Research Consortium of Canada (SPARCC) score of the sacroiliac (SI) joints. METHODS: Patients with back pain (≥ 3 months, ≤ 2 years, onset < 45 years) were included in the SPACE cohort (SpondyloArthritis Caught Early). Patients with (possible) axial spondyloarthritis had followup visits after 3 and 12 months and were treated according to clinical practice. Magnetic resonance imaging (MRI) of the SI joints (MRI-SI) was scored in 2 independent campaigns (campaign 1: at baseline and 3 months; campaign 2: at baseline, 3 months, and 12 months) by 2 different blinded reader pairs, applying the Assessment of Spondyloarthritis International Society (ASAS) definition (MRI-SI+ vs MRI-SI-; discordant cases were adjudicated by a third reader) and SPARCC score (mean of 2 agreeing readers). Calculations were made for agreement between SPARCC score cutoff values and a consensus judgment of MRI-SI+ (ASAS definition) as external standard, change in SPARCC score, and smallest detectable changes (SDC) over 3 and 12 months. RESULTS: SPARCC score ≥ 2 showed best agreement with MRI-SI+ in both campaigns. Regarding observed changes in relation to SDC, SPARCC score changed in 70/151 patients; 26/70 patients changed > SDC (3.4), of whom 20 patients received stable treatment over 3 months in campaign 1. Over 3 months, 20/68 patients showed changes in SPARCC score; 11/20 > SDC (2.1), of whom 8 patients received stable treatment. Over 1 year, 23/74 patients changed their SPARCC score; 14/23 changed > SDC (2.4), of whom 7 received stable treatment in campaign 2. CONCLUSION: SPARCC score ≥ 2 can be used as surrogate for a consensus judgment of MRI-SI+ (ASAS definition) in clinical trials. The SDC ranged from 2.1-3.4 dependent on reader pair and were close to the proposed minimum important change of 2.5.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.009 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".