AB1179 VARIATION OF INTER-RATER RELIABILITY FOR BONE MARROW LESION SCORES IN DIFFERENT REGIONS OF THE KNEE: A STUDY USING KIMRISS
Bibliographic record
Abstract
Background Semi-quantitative magnetic resonance imaging (MRI) scoring is more challenging in some anatomic regions than others. Bone marrow lesion (BML) is an important scoring feature and may be more closely related to clinical outcomes in arthritis in some anatomic regions than others. The Knee Inflammation MRI Scoring System (KIMRISS) scores BML and synovitis-effusion from fluid-sensitive sagittal MRI[1]. Via web-based interface, readers place ready-made overlays onto scorable portions of each bone, divided into grid squares which are scored dichotomously (yes/no) for presence of BML. The total KIMRISS BML score, which sums positive grid squares on all slices, has been previously shown to have high inter-rater reliability for both status and change. Objectives We aimed to evaluate patterns of regional variation in inter-rater reliability for BML across the knee via KIMRISS grid elements. This granular analysis serves as a foundation for further study of the relationship of precise BML location to clinical outcomes. Methods Baseline MRI from 61 patients in the Osteoarthritis Initiative (OAI) dataset were scored by 8 trained readers using KIMRISS scoring. Total scores for each of the 28 grid overlay regions (14 femur, 10 tibia, 4 patella) were calculated by adding up positive scores from all scorable MRI slices. We calculated descriptive statistics and interclass correlation coefficients (ICC) for all 28 reader pairs, then combined any neighbouring grid locations with ICC <0.5 before recalculating the ICC. Results The 8-reader mean (SD) BML scores at each scoring location (Table 1) showed that the mean reader pair ICC was very good (>=0.80) or good (0.70-0.80) for 17/28 (61%) of the grid regions (Figure 1). The lowest ICCs occurred at the inferior patella [P4; mean (SD) ICC=0.52 (0.19)], a small interior region of the femur [mean (SD) ICC=0.49 (0.26)], and the two most posterior regions of the tibia [mean (SD) ICC=0.41 (0.26) and 0.31 (0.33)]. Lower reliability was also found in regions of tibia not immediately adjacent to the tibiofemoral joint (ICC range: 0.60-0.67) as well as posterior non-articular regions of the femur (ICC range: 0.59-0.65). Combining the posterior tibial regions scores (T5+T55) slightly improved ICC compared to separate scores [mean (SD) ICC= 0.40 (0.29)]. Conclusion Most individual KIMRISS grid squares were scored with high reliability. Decreased reliability in certain regions may be due to lower incidence of BML or anatomical features that lead to difficulty of interpretation at specific locations. Combining neighbouring regions with lower reliability, such as those at the posterior tibia, may increase reliability if discrepancy arises from assigning a positive score for the same lesion to different grid squares between readers. The patterns of reliability demonstrated here support further study of the impact of precise BML location status scores on clinical outcomes. Reference [1]Jaremko JL, Jeffery D, Buller M, et al Preliminary validation of the Knee Inflammation MRI Scoring System (KIMRISS) for grading bone marrow lesions in osteoarthritis of the knee: data from the Osteoarthritis Initiative RMD Open 2017;3:e000355. doi: 10.1136/rmdopen-2016-000355 Acknowledgements: NIL. Disclosure of Interests Stephanie Wichuk: None declared, Paul Bird Consultant of: AbbVie, Bristol Myers Squibb, Celgene, Janssen, MSD, Novartis, Pfizer, Roche, UCB, Iris Eshed: None declared, Robert G Lambert Consultant of: Calyx, CARE Arthritis, Image Analysis Group (IAG), Walter P Maksymowych Consultant of: AbbVie, Bristol Myers Squibb, Boehringer, Celgene, Eli Lilly, Galapagos, Janssen, Novartis, Pfizer, and UCB, Grant/research support from: AbbVie, Novartis, Pfizer, and UCB, Andrew McReynolds: None declared, Susanne Juhl Pedersen Speakers bureau: MSD, Pfizer, AbbVie, UCB, Novartis, Consultant of: AbbVie, UCB, Novartis, Grant/research support from: AbbVie, MSD, and Novartis, Ulrich Weber: None declared, Joel Paschke: None declared, Jacob L Jaremko: None declared.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.025 | 0.047 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".