MétaCan
Menu
← Back to cohort

AB1179 VARIATION OF INTER-RATER RELIABILITY FOR BONE MARROW LESION SCORES IN DIFFERENT REGIONS OF THE KNEE: A STUDY USING KIMRISS

2023· article· en· W4379650509 on OpenAlexaff
Stephanie Wichuk, Paul Bird, Iris Eshed, R. Lambert, Walter P. Maksymowych, Andrew McReynolds, Susanne Juhl Pedersen, Ulrich Weber, Joel Paschke, Jacob L. Jaremko

Bibliographic record

Venuenot available
Typearticle
Languageen
FieldMedicine
TopicOsteoarthritis Treatment and Mechanisms
Canadian institutionsUniversity of Alberta
Fundersnot available
KeywordsMedicineMagnetic resonance imagingOsteoarthritisReliability (semiconductor)Sagittal planeInterclass correlationSynovitisIntraclass correlationRank correlationRadiologyNuclear medicineArthritisPathologyStatisticsMathematicsInternal medicine

Abstract

fetched live from OpenAlex

Background Semi-quantitative magnetic resonance imaging (MRI) scoring is more challenging in some anatomic regions than others. Bone marrow lesion (BML) is an important scoring feature and may be more closely related to clinical outcomes in arthritis in some anatomic regions than others. The Knee Inflammation MRI Scoring System (KIMRISS) scores BML and synovitis-effusion from fluid-sensitive sagittal MRI[1]. Via web-based interface, readers place ready-made overlays onto scorable portions of each bone, divided into grid squares which are scored dichotomously (yes/no) for presence of BML. The total KIMRISS BML score, which sums positive grid squares on all slices, has been previously shown to have high inter-rater reliability for both status and change. Objectives We aimed to evaluate patterns of regional variation in inter-rater reliability for BML across the knee via KIMRISS grid elements. This granular analysis serves as a foundation for further study of the relationship of precise BML location to clinical outcomes. Methods Baseline MRI from 61 patients in the Osteoarthritis Initiative (OAI) dataset were scored by 8 trained readers using KIMRISS scoring. Total scores for each of the 28 grid overlay regions (14 femur, 10 tibia, 4 patella) were calculated by adding up positive scores from all scorable MRI slices. We calculated descriptive statistics and interclass correlation coefficients (ICC) for all 28 reader pairs, then combined any neighbouring grid locations with ICC <0.5 before recalculating the ICC. Results The 8-reader mean (SD) BML scores at each scoring location (Table 1) showed that the mean reader pair ICC was very good (>=0.80) or good (0.70-0.80) for 17/28 (61%) of the grid regions (Figure 1). The lowest ICCs occurred at the inferior patella [P4; mean (SD) ICC=0.52 (0.19)], a small interior region of the femur [mean (SD) ICC=0.49 (0.26)], and the two most posterior regions of the tibia [mean (SD) ICC=0.41 (0.26) and 0.31 (0.33)]. Lower reliability was also found in regions of tibia not immediately adjacent to the tibiofemoral joint (ICC range: 0.60-0.67) as well as posterior non-articular regions of the femur (ICC range: 0.59-0.65). Combining the posterior tibial regions scores (T5+T55) slightly improved ICC compared to separate scores [mean (SD) ICC= 0.40 (0.29)]. Conclusion Most individual KIMRISS grid squares were scored with high reliability. Decreased reliability in certain regions may be due to lower incidence of BML or anatomical features that lead to difficulty of interpretation at specific locations. Combining neighbouring regions with lower reliability, such as those at the posterior tibia, may increase reliability if discrepancy arises from assigning a positive score for the same lesion to different grid squares between readers. The patterns of reliability demonstrated here support further study of the impact of precise BML location status scores on clinical outcomes. Reference [1]Jaremko JL, Jeffery D, Buller M, et al Preliminary validation of the Knee Inflammation MRI Scoring System (KIMRISS) for grading bone marrow lesions in osteoarthritis of the knee: data from the Osteoarthritis Initiative RMD Open 2017;3:e000355. doi: 10.1136/rmdopen-2016-000355 Acknowledgements: NIL. Disclosure of Interests Stephanie Wichuk: None declared, Paul Bird Consultant of: AbbVie, Bristol Myers Squibb, Celgene, Janssen, MSD, Novartis, Pfizer, Roche, UCB, Iris Eshed: None declared, Robert G Lambert Consultant of: Calyx, CARE Arthritis, Image Analysis Group (IAG), Walter P Maksymowych Consultant of: AbbVie, Bristol Myers Squibb, Boehringer, Celgene, Eli Lilly, Galapagos, Janssen, Novartis, Pfizer, and UCB, Grant/research support from: AbbVie, Novartis, Pfizer, and UCB, Andrew McReynolds: None declared, Susanne Juhl Pedersen Speakers bureau: MSD, Pfizer, AbbVie, UCB, Novartis, Consultant of: AbbVie, UCB, Novartis, Grant/research support from: AbbVie, MSD, and Novartis, Ulrich Weber: None declared, Joel Paschke: None declared, Jacob L Jaremko: None declared.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.025
metaresearch head score (Gemma)0.047
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.025
Threshold uncertainty score0.134

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0250.047
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.002
Science and technology studies0.0010.001
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.050
GPT teacher head0.312
Teacher spread0.262 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2023
Admission routes1
Has abstractyes

Explore more

Same topicOsteoarthritis Treatment and Mechanisms→French-language works237,207→