Canada-Denmark MRI scoring system of the spine in patients with axial spondyloarthritis: updated definitions, scoring rules and inter-reader reliability in a multiple reader setting
Bibliographic record
Abstract
Objective: To validate the Canada-Denmark (CANDEN) MRI scoring system for the spine in axial spondyloarthritis with updated lesion definitions. Methods: Lesion definitions in the CANDEN system were updated and illustrated by a consensus set of reference images. Sagittal spine MRIs of 40 patients with axial spondyloarthritis obtained at baseline and at week 52 after initiation of treatment with the tumour necrosis factor inhibitor golimumab were evaluated in unknown chronology by seven readers blinded to all other data. Results: CANDEN MRI spine inflammation score had very good reliability for status scores (single-measure intraclass correlation coefficient (ICC) of 21 reader pairs median of 0.91 (IQR 0.88-0.92)) and change scores (ICC 0.88 (0.86-0.92)). CANDEN MRI spine fat score had good to very good reliability for status scores (ICC 0.79 (0.75-0.86)) and moderate to good reliability for detecting change (ICC 0.59 (0.46-0.73)). CANDEN MRI spine bone erosion score and CANDEN MRI spine new bone formation score had slight to moderate reliability for status scores (ICC 0.38 (0.32-0.52) and 0.39 (0.27-0.49), respectively). Conclusion: The CANDEN MRI spine scoring system allows a comprehensive evaluation of inflammation, fat, bone erosion and new bone formation of the spine in patients with axial spondyloarthritis. It demonstrated very good reliability for detecting change in inflammation, moderate to good reliability for detecting change in fat, and slight to moderate reliability for detecting bone erosions and new bone formation. Studies with longer follow-up or patients with more advanced spinal involvement may be needed to reliably detect change in bone erosion and new bone formation scores. Trial registration number: NCT02011386.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.028 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".