Re-Evaluating the Impact of Including Patients with Bilateral Conditions in Orthopaedic Clinical Research Studies
Bibliographic record
Abstract
BACKGROUND: Orthopaedic studies frequently include subjects with bilateral conditions. Failure to account for bilateral conditions can lead to spurious associations. The performance of different methods for addressing this issue, especially in populations that include subjects with unilateral and bilateral conditions, has not been rigorously evaluated. The purpose of the present study was to test 3 different methods for analyzing bilateral data: (1) analyzing all limbs as independent subjects (naïve), (2) randomly selecting 1 limb per subject (random), and (3) accounting for correlation between limbs with use of a linear mixed model (LMM). METHODS: We simulated a hypothetical randomized controlled trial in which Western Ontario and McMaster Universities Osteoarthritis Index (WOMAC) scores were collected at a baseline and a 2-year visit. We simulated 2 scenarios: Scenario 1 (in which there was truly no difference between groups [mean difference = 0]) and Scenario 2 (in which there was truly a difference between groups [mean difference = 10]). We varied the prevalence of bilateral involvement from 10% to 100% within each scenario. We evaluated method performance on the basis of bias (difference from the simulated true effect), power (1 - type-II error), type-1 error rate, and 95% confidence interval (CI) coverage. RESULTS: Bias (difference from simulated true effect) was similar across all methods. In Scenario 2 (true difference between groups), CI coverage was lowest with use of the naïve method (median, 87.8%; range, 85.3% to 93.5%) relative to the random method (median, 95.1%; range, 94.5% to 95.6%) and the LMM method (median, 95.1%; range, 94.5% to 95.5%). In Scenario 1 (no difference between groups), the type-1 error rate was highest for the naïve method (median, 11.3%; range, 6.7% to 14.7%) relative to the LMM method (median, 4.9%; range, 4.5% to 5.3%) and the random method (median, 5.0%; range, 4.5% to 5.2%). CONCLUSIONS: Failure to account for bilateral conditions led to biased CIs and an increased type-1 error rate. Due to the fact that bias was similar across the methods, decreased model performance using the naïve method was likely attributable to underestimation of the standard error. Orthopaedic studies involving subjects with bilateral conditions warrant special considerations that can be addressed using simple (random) or more complex (LMM) methods. CLINICAL RELEVANCE: Adherence to robust methodological practices is an essential but underappreciated component of the translation of evidence into clinical practice. Our work is meant to be educational, providing clinical researchers with the knowledge and skills to address a common challenge within the field.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.644 | 0.833 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.005 | 0.010 |
| Bibliometrics | 0.004 | 0.005 |
| Science and technology studies | 0.002 | 0.005 |
| Scholarly communication | 0.007 | 0.006 |
| Open science | 0.005 | 0.006 |
| Research integrity | 0.005 | 0.004 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".