Scoring magnetic resonance imaging (MRI) inflammation and structural lesions in sacroiliac joints of patients with axial spondyloarthritis: assessment of all MRI slices of the cartilaginous compartment versus standardized six or five slices
Bibliographic record
Abstract
Objectives: The Spondyloarthritis Research Consortium of Canada (SPARCC) sacroiliac joint (SIJ) scoring system assesses six or five (6/5) semicoronal magnetic resonance imaging (MRI) slices for inflammation/structural lesions in patients with axial spondyloarthritis (axSpA). However, the cartilaginous SIJ compartment may be visible in a few additional slices. The objective was to investigate interreader reliability, sensitivity to change, and classification of MRI scans as positive or negative for various lesion types using an ‘all slices’ approach versus standard SPARCC scoring of 6/5 slices.Method: Fifty-three axSpA patients were treated with the tumour necrosis factor inhibitor golimumab and followed with serial MRI scans at weeks 0, 4, 16, and 52. The most anterior and posterior slices covering the cartilaginous compartment and the transitional slice were identified. Scores for inflammation, fat metaplasia, erosion, backfill, and ankylosis in the cartilaginous SIJ compartment were calculated for the ‘all slices’ approach and the 6/5 slices standard.Results: By the ‘all slices’ approach, three readers scored mean 7.2, 7.7, and 7.0 slices per MRI scan. Baseline and change scores for the various lesion types closely correlated between the two approaches (Pearson’s rho ≥ 0.95). Inflammation score was median 13 (interquartile range 6–21, range 0–49) for 6/5 slices versus 14 (interquartile range 6–23, range 0–69) for all slices at baseline. Interreader reliability, sensitivity to change, and classification of MRI scans as positive or negative for various lesion types were similar.Conclusion: The standardized 6/5 slices approach showed no relevant differences from the ‘all slices’ approach and, therefore, is equally suited for monitoring purposes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".