Lack of standardization among clinical trials of injection therapies for knee osteoarthritis: a systematic review
Bibliographic record
Abstract
Purpose: Osteoarthritis (OA) of the knee is a debilitating, expensive, and prevalent disease, and interest in the non-surgical management of knee OA has grown recently. Our objective was to systematically assess the level of heterogeneity among all clinical trials and published studies regarding injections for knee osteoarthritis, in terms of treatment of interest, outcomes evaluated, and time points of outcome assessment.Methods: The Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines were utilized to review all published studies and publically available clinical trials from 1 January 2013 to 3 May 2019evaluating intra-articular injections to treat knee OA. Their treatment group and specifics of methodology were scrutinized and compared.Results: 84 published studies and 114 clinical trials were included. Within the 84 published studies, the most common injection treatment studied was hyaluronic acid [N = 22; 26.2%]. In total, 29 different injection treatment groups were utilized. The most common time point for patient evaluation post-injection was 6 months (N = 33 studies; 50.0%), and ranged from 1 week (N = 9 studies; 13.6%) to 7 years (N = 1 study; 1.5%). The most common patient-reported outcome (PRO) measure assessed in the included studies was Western Ontario and McMaster’s University Osteoarthritis Index (WOMAC) [N = 44 studies; 66.7%]. For the 114 clinical trials identified, the most common injection treatment studied is platelet-rich plasma in isolation (N = 19; 16.7%). Forty-two different injection treatment types/groups are utilized. The most common PRO measure assessed was WOMAC (N = 77 trials; 67.5%). Overall there were 34 different patient-reported outcome measures used.Conclusions: Research efforts to find the most effective injection therapy for knee OA continue with a tremendous number of injection therapies still being evaluated. Substantial heterogeneity exists in these completed and ongoing trials in terms of patient demographics, OA grades, outcome scores and relatively short-term timing of assessments, with no clear standardization of testing protocol despite proposing to answer the same clinical question. We recommend that studies of this genre going forward be standardized in terms of outcome measures and longer-term follow-up time points, and should incorporate functional assessment evaluations and imaging studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.185 | 0.449 |
| Meta-epidemiology (narrow) | 0.002 | 0.003 |
| Meta-epidemiology (broad) | 0.017 | 0.020 |
| Bibliometrics | 0.015 | 0.014 |
| Science and technology studies | 0.001 | 0.004 |
| Scholarly communication | 0.007 | 0.006 |
| Open science | 0.004 | 0.004 |
| Research integrity | 0.004 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".