Reliability of Magnetic Resonance Imaging Interpretation of Lateral Discoid Meniscus: A Multicenter Study
Bibliographic record
Abstract
Background: Lateral discoid meniscus (LDM) often present with complex morphology that can be challenging to assess and treat arthroscopically. Preoperative magnetic resonance imaging (MRI) is frequently used for diagnosis and surgical planning, however it is not known whether surgeons are reliable and consistent in their interpretation of MRI findings. Hypothesis/Purpose: We hypothesize that surgeons experienced in treating LDM are able to reliably interpret discoid pathology using MRI, with the exception of evaluating dynamic factors (i.e. instability). Methods: Forty five discoid meniscus MRI selected from a pool of surgical discoid meniscus cases were included in this review. Five reviewers (experienced pediatric sports medicine surgeons) who were not involved in the MRI selection process performed independent review of each MRI to determine discoid meniscus classification. More than 4 weeks later, a second reading was performed by 3 of the 5 reviewers. Interobserver and intraobserver reliability of the primary (width, presence of instability or tear) and secondary (location of instability or tear, tear type) rating factors was assessed using the Fleiss κ coefficient, designed for multiple readers with nominal variables (reliability: fair, 0.21–0.40; moderate, 0.41–0.60; substantial, 0.61–0.80; excellent, 0.81–1.00). Reliability is reported as κ (95% CI). Results: Interobserver reliability of assessment of meniscal width was substantial 0.67 (0.58-0.76), and intraobserver was moderate 0.52 (0.35-0.69). Assessment of presence of peripheral instability had moderate interobserver reliability 0.54 (0.45-0.63) and substantial intraobserver 0.61 (0.44-0.78). Presence of tear had fair interobserver reliability 0.39 (0.29–0.48) and substantial intraobserver 0.68 (0.51-0.85). When identifying location of instability or tear, interobserver reliability was moderate while intraobserver agreement was substantial for assessment of anterior instability (0.47 (0.37-0.56) and 0.68 (0.51-0.85), respectively); posterior instability (0.56 (0.47-0.65) and 0.62 (0.45-0.79), respectively); and posterior tear (0.41 (0.32-0.50) and 0.69 (0.52-0.86), respectively). The exception was identifying anterior tear, with fair interobserver reliability 0.33 (0.23-0.42) and moderate intraobserver 0.56 (0.39-0.73); Interobserver reliability was fair for tear type 0.34 (0.28-0.41) and intraobserver reliability moderate 0.55 (0.43-0.67). Conclusion: Orthopaedic surgeons experienced in the treatment of LDM vary from each other in their classification using MRI, especially with regard to assessment of meniscus tears. MRI evaluation may be helpful to diagnose discoid by width and identify presence of instability, two major factors in the decision to proceed with surgery. However, definitive treatment should be guided by a comprehensive arthroscopic evaluation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.035 | 0.089 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".