A Historical Analysis of Randomized Controlled Trials in Anterior Cruciate Ligament Surgery
Bibliographic record
Abstract
BACKGROUND: The purpose of this systematic review was to comprehensively assess the quality of reporting of randomized controlled trials (RCTs) relating to anterior cruciate ligament (ACL) reconstruction. Specifically, this review explored factors related to the quality of the RCTs and trends in the quality of reporting over time. METHODS: The online databases PubMed, Ovid (MEDLINE), and Embase were used to search for all RCTs on the topic of ACL reconstruction from database inception until April 14, 2016. The quality of reporting was evaluated using the Detsky quality index and the Consolidated Standards of Reporting Trials (CONSORT) checklist for reporting trials of nonpharmacologic treatments. A multivariate regression analysis was used to assess predictors of quality reporting. RESULTS: The online search yielded 2,933 articles, 412 of which met the inclusion criteria and were assessed for quality of reporting. There was a significant (p < 0.0001) increase in the number of RCTs published over time. The mean Detsky score (and standard deviation) across all included RCTs was 68.9% ± 13.2%. The strongest predictors of quality reporting were the inclusion of a CONSORT flow diagram (β-coefficient, 10.0; 95% confidence interval [CI]: 8.45 to 11.61; p < 0.0001) and being published in the year 2009 or later (β-coefficient, 5.2; 95% CI: 3.87 to 6.45; p < 0.0001). The factors demonstrating the greatest improvement over time were the inclusion of a full description of the randomization procedure (p = 0.001) and prospective calculation of the sample size (p = 0.002). CONCLUSIONS: There has been a significant increase in both the quantity and quality of RCTs relating to ACL reconstruction over time. Specifically, the reporting of a methodologically sound randomization process and prospective calculation of sample size have significantly improved in recent years. However, since the year 2009, the number of trials and reporting in these trials has remained relatively consistent. The use of a CONSORT flow diagram is a strong predictor of high-quality reporting.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | MetaresearchMeta-epidemiology (broad) Domain: Reporting · Genre: Review About the Canadian research system: no · About a Canadian topic: no | Systematic review | high |
| gpt | MetaresearchMeta-epidemiology (broad)Bibliometrics Domain: Reporting · Genre: Review About the Canadian research system: no · About a Canadian topic: no | Systematic review | high |
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.080 | 0.033 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.120 | 0.042 |
| Bibliometrics | 0.011 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
The models disagree on parts of this classification; every voice is preserved in the section at the end of the page.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".