Exploration of heterogeneity of treatment effects across exercise-based interventions for knee osteoarthritis
Bibliographic record
Abstract
Objective: Variability exists in the degree of improvement patients experience following exercise-based interventions (EBIs) for knee osteoarthritis (KOA), but understanding of this heterogeneity is limited. Using a machine learning approach, this study leveraged data from two randomized controlled trials (RCTs) to identify patient characteristics contributing to differential treatment effects. Design: The RCTs enrolled n = 621 patients and evaluated three EBIs (group-based physical therapy (PT), individual PT, and a Stepped Exercise Program) and an education control group. The primary outcome was change in total Western Ontario and McMaster Universities Osteoarthritis Index (WOMAC) score from baseline to end of treatment. Predictors included 25 demographic, clinical, and psychosocial characteristics. Three metalearners with three machine learning algorithms each and a simple interpretable model-based regression tree were used to identify subgroups with differential treatment effects. Fit was evaluated with holdout/validation data using root mean square error and mean absolute error. Results: The regression tree model outperformed all 9 metalearner models. Tree results suggested group-based PT yielded the largest improvement in mean WOMAC score. Only two subgroups were identified: baseline WOMAC score≤44 versus >44. Group-based PT was the optimal treatment regardless of baseline WOMAC score, but results were more ambiguous for patients with higher initial WOMAC score. For all 3 EBIs, patients with higher baseline WOMAC score made greater improvements. Conclusion: Results suggest individuals with moderate or greater KOA symptoms may benefit more from EBIs than those with less severe symptoms and that group-based PT is a promising approach for KOA.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".