Improving the specification of the target difference in the sample size calculation of a randomised trial of treatments for osteoarthritis
Bibliographic record
Abstract
The sample size of a clinical trial is the number of participants the trial aims to recruit. Sample size is a critical aspect of clinical trial design and has ethical and financial implications. The sample size depends on the target difference, the difference in outcome that the trial is powered to detect. This thesis aims to improve methods for specifying the target difference in randomised trials of osteoarthritis. I conducted a systematic review of sample size calculations in hip and knee osteoarthritis trials published in 2016. It found that most sample size calculations were poorly reported and could not be reproduced. The target difference in the sample size calculation was commonly justified by a published minimum clinically important difference (MCID). Several versions of the WOMAC (Western Ontario and McMaster Universities Osteoarthritis Index) were commonly used in hip and knee osteoarthritis trials. It was often unclear which version was used, hindering interpretation of trial results. I conducted a discrete choice experiment examining patient preferences when choosing between osteoarthritis medications. Duration of treatment effect was shown to be important to participants, viewed with similar importance to the amount of symptom relief provided and risks of the treatment. I analysed a cohort of people with osteoarthritis and showed that MCID estimates for the WOMAC varied across different follow-up time points. However, there was no visual trend in the change in MCID estimates over time. Longitudinal methods were feasible to calculate MCID estimates, but did not improve precision. A simulation study that I conducted found that the pattern of the treatment effect (its duration and consistency) affected the optimal statistical method of analysis for a randomised trial using the WOMAC as the primary outcome. Future research is needed to examine whether the findings are generalisable to different datasets, outcome measures and health conditions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.027 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".