Application of minimal important differences in degenerative knee disease outcomes: a systematic review and case study to inform <i>BMJ</i> Rapid Recommendations
Bibliographic record
Abstract
Objectives To identify the most credible anchor-based minimal important differences (MIDs) for patient important outcomes in patients with degenerative knee disease, and to inform BMJ Rapid Recommendations for arthroscopic surgery versus conservative management Design Systematic review. Outcome measures Estimates of anchor-based MIDs, and their credibility, for knee symptoms and health-related quality of life (HRQoL). Data sources MEDLINE, EMBASE and PsycINFO. Eligibility criteria We included original studies documenting the development of anchor-based MIDs for patient-reported outcomes (PROs) reported in randomised controlled trials included in the linked systematic review and meta-analysis and judged by the parallel BMJ Rapid Recommendations panel as critically important for informing their recommendation: measures of pain, function and HRQoL. Results 13 studies reported 95 empirically estimated anchor-based MIDs for 8 PRO instruments and/or their subdomains that measure knee pain, function or HRQoL. All studies used a transition rating (global rating of change) as the anchor to ascertain the MID. Among PROs with more than 1 estimated MID, we found wide variation in MID values. Many studies suffered from serious methodological limitations. We identified the following most credible MIDs: Western Ontario and McMaster University Osteoarthritis Index (WOMAC; pain: 12, function: 13), Knee injury and Osteoarthritis Outcome Score (KOOS; pain: 12, activities of daily living: 8) and EuroQol five dimensions Questionnaire (EQ-5D; 0.15). Conclusions We were able to distinguish between more and less credible MID estimates and provide best estimates for key instruments that informed evidence presentation in the associated systematic review and judgements made by the Rapid Recommendation panel. Trial registration number CRD42016047912.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.006 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".