Pilot randomized controlled trials in the orthopaedic surgery literature: a systematic review
Bibliographic record
Abstract
BACKGROUND: The primary objective of this systematic review is to examine the characteristics of pilot randomized controlled trials (RCTs) in the orthopaedic surgery literature, including the proportion framed as feasibility trials and those that lead to definitive RCTs. This review aim to answer the question of whether pilot RCTs lead to definitive RCTs, whilst investigating the quality, feasibility and overall publication trends of orthopaedic pilot trials. METHODS: Pilot RCTs in the orthopaedic literature were identified from three electronic databases (EMBASE, MEDLINE, and Pubmed) searched from database inception to January 2018. Search criteria included the evaluation of at least one orthopaedic surgical intervention, research on humans, and publication in English. Two reviewers independently screened the pool of pilot trials, and conducted a search for corresponding definitive trials. Screened pilot RCTs were assessed for feasibility outcomes related to efficiency, cost, and/or timeliness of a large-scale clinical trial involving a surgical intervention. The quality of the pilot and definitive trials were assessed using the Checklist to Evaluate a Report of a Non-Pharmacological Trial (CLEAR NPT). RESULTS: The initial search for pilot RCTs yielded 3857 titles, of which 49 articles were relevant for this review. 73.5% (36/49) of the orthopaedic pilot RCTs were framed as feasibility trials. Of these, 5 corresponding definitive trials (10.2%) were found, of which four were published and one ongoing. Based on author responses, the lack of a definitive RCT following the pilot trial was attributed to a lack of funding, inadequacies in recruitment, and belief that the pilot RCT sufficiently answered the research question. CONCLUSIONS: Based on this systematic review, most pilot RCTs were characterized as feasibility trials. However, the majority of published pilot RCTs did not lead to definitive trials. This discrepancy was mainly attributed to poor feasibility (e.g. poor recruitment) and lack of funding for an orthopaedic surgical definitive trial. In recent years this discrepancy may be due to researchers saving on time and cost by rolling their pilot patients into the definitive RCT rather than publish a separate pilot trial.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | MetaresearchMeta-epidemiology (broad) Domain: Methods · Genre: Review About the Canadian research system: no · About a Canadian topic: no | Systematic review | high |
| gpt | Metaresearch Domain: Methods · Genre: Review About the Canadian research system: no · About a Canadian topic: no | Systematic review | high |
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.870 | 0.890 |
| Meta-epidemiology (narrow) | 0.003 | 0.001 |
| Meta-epidemiology (broad) | 0.197 | 0.203 |
| Bibliometrics | 0.003 | 0.009 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.003 | 0.000 |
| Open science | 0.007 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
The models disagree on parts of this classification; every voice is preserved in the section at the end of the page.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".