Magnetic Resonance Imaging (MRI) of the Knee as an Outcome Measure in Juvenile Idiopathic Arthritis: An OMERACT Reliability Study on MRI Scales
Bibliographic record
Abstract
OBJECTIVE: There is increasing evidence that early therapeutic intervention improves longterm joint outcome in juvenile idiopathic arthritis (JIA). Given the existence of highly effective treatments, there is an urgent need for reliable and accurate measures of disease activity and joint damage in JIA. Our objective was to assess the reliability of 2 magnetic resonance imaging (MRI) scoring methods: the Juvenile Arthritis MRI Scoring (JAMRIS) system and the International Prophylaxis Study Group (IPSG) consensus score, for evaluating disease status of the knee in patients with JIA. METHODS: Four international readers independently scored an MRI dataset of 25 JIA patients with clinical knee involvement. Synovial thickening, joint effusion, bone marrow changes, cartilage lesions, bone erosions, and subchondral cysts were scored using the JAMRIS and IPSG systems. Further, synovial enhancement, infrapatellar fat pad heterogeneity, tendinopathy, and enthesopathy were scored. Interreader reliability was analyzed by using the generalized κ, ICC, and the smallest detectable difference (SDD). RESULTS: ICC regarding interreader reliability ranged from 0.33 (95% CI 0.12-0.52, SDD = 0.29) for enthesopathy up to 0.95 (95% CI 0.92-0.97, SDD = 3.19) for synovial thickening. Good interreader reliability was found concerning joint effusion (ICC 0.93, 95% CI 0.89-0.95, SDD = 0.51), synovial enhancement (ICC 0.90, 95% CI 0.85-0.94, SDD = 9.85), and bone marrow changes (ICC 0.87, 95% CI 0.80-0.92, SDD = 10.94). Moderate to substantial reliability was found concerning cartilage lesions and bone erosions (ICC 0.55-0.72, SDD 1.41-13.65). CONCLUSION: The preliminary results are promising for most of the scored JAMRIS and IPSG items. However, further refinement of the scoring system is warranted for unsatisfactorily reliable items such as bone erosions, cartilage lesions, and enthesopathy.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".