AB1179 VARIATION OF INTER-RATER RELIABILITY FOR BONE MARROW LESION SCORES IN DIFFERENT REGIONS OF THE KNEE: A STUDY USING KIMRISS
Notice bibliographique
Résumé
Background Semi-quantitative magnetic resonance imaging (MRI) scoring is more challenging in some anatomic regions than others. Bone marrow lesion (BML) is an important scoring feature and may be more closely related to clinical outcomes in arthritis in some anatomic regions than others. The Knee Inflammation MRI Scoring System (KIMRISS) scores BML and synovitis-effusion from fluid-sensitive sagittal MRI[1]. Via web-based interface, readers place ready-made overlays onto scorable portions of each bone, divided into grid squares which are scored dichotomously (yes/no) for presence of BML. The total KIMRISS BML score, which sums positive grid squares on all slices, has been previously shown to have high inter-rater reliability for both status and change. Objectives We aimed to evaluate patterns of regional variation in inter-rater reliability for BML across the knee via KIMRISS grid elements. This granular analysis serves as a foundation for further study of the relationship of precise BML location to clinical outcomes. Methods Baseline MRI from 61 patients in the Osteoarthritis Initiative (OAI) dataset were scored by 8 trained readers using KIMRISS scoring. Total scores for each of the 28 grid overlay regions (14 femur, 10 tibia, 4 patella) were calculated by adding up positive scores from all scorable MRI slices. We calculated descriptive statistics and interclass correlation coefficients (ICC) for all 28 reader pairs, then combined any neighbouring grid locations with ICC <0.5 before recalculating the ICC. Results The 8-reader mean (SD) BML scores at each scoring location (Table 1) showed that the mean reader pair ICC was very good (>=0.80) or good (0.70-0.80) for 17/28 (61%) of the grid regions (Figure 1). The lowest ICCs occurred at the inferior patella [P4; mean (SD) ICC=0.52 (0.19)], a small interior region of the femur [mean (SD) ICC=0.49 (0.26)], and the two most posterior regions of the tibia [mean (SD) ICC=0.41 (0.26) and 0.31 (0.33)]. Lower reliability was also found in regions of tibia not immediately adjacent to the tibiofemoral joint (ICC range: 0.60-0.67) as well as posterior non-articular regions of the femur (ICC range: 0.59-0.65). Combining the posterior tibial regions scores (T5+T55) slightly improved ICC compared to separate scores [mean (SD) ICC= 0.40 (0.29)]. Conclusion Most individual KIMRISS grid squares were scored with high reliability. Decreased reliability in certain regions may be due to lower incidence of BML or anatomical features that lead to difficulty of interpretation at specific locations. Combining neighbouring regions with lower reliability, such as those at the posterior tibia, may increase reliability if discrepancy arises from assigning a positive score for the same lesion to different grid squares between readers. The patterns of reliability demonstrated here support further study of the impact of precise BML location status scores on clinical outcomes. Reference [1]Jaremko JL, Jeffery D, Buller M, et al Preliminary validation of the Knee Inflammation MRI Scoring System (KIMRISS) for grading bone marrow lesions in osteoarthritis of the knee: data from the Osteoarthritis Initiative RMD Open 2017;3:e000355. doi: 10.1136/rmdopen-2016-000355 Acknowledgements: NIL. Disclosure of Interests Stephanie Wichuk: None declared, Paul Bird Consultant of: AbbVie, Bristol Myers Squibb, Celgene, Janssen, MSD, Novartis, Pfizer, Roche, UCB, Iris Eshed: None declared, Robert G Lambert Consultant of: Calyx, CARE Arthritis, Image Analysis Group (IAG), Walter P Maksymowych Consultant of: AbbVie, Bristol Myers Squibb, Boehringer, Celgene, Eli Lilly, Galapagos, Janssen, Novartis, Pfizer, and UCB, Grant/research support from: AbbVie, Novartis, Pfizer, and UCB, Andrew McReynolds: None declared, Susanne Juhl Pedersen Speakers bureau: MSD, Pfizer, AbbVie, UCB, Novartis, Consultant of: AbbVie, UCB, Novartis, Grant/research support from: AbbVie, MSD, and Novartis, Ulrich Weber: None declared, Joel Paschke: None declared, Jacob L Jaremko: None declared.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,025 | 0,047 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,002 | 0,002 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,001 | 0,001 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,002 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».