Have We Progressed in the Surgical Literature? Thirty-Year Trends in Clinical Studies in 3 Surgical Journals
Bibliographic record
Abstract
BACKGROUND: We practice in an era of evidence-based medicine. In 1993, Solomon and McLeod published an article examining study designs in 3 surgical journals from 1980 and 1990. OBJECTIVE: The purpose of this study was to evaluate subsequent 30-year trends in the quality of selected literature. DESIGN: All of the articles from Diseases of the Colon & Rectum, Surgery, and the British Journal of Surgery during 2000 and 2010 were classified by study design. Nonclinical studies were substratified by animal/laboratory, surgical technique, editorial/review, or miscellaneous articles. Clinical articles were categorized as case or comparative studies, further categorized by study design, and rated on a 10-point scale to determine strength. We compared interobserver reliability using a random sample. SETTING: This study was conducted at 3 North American medical centers. PATIENTS: Patients described in the scope of the literature were included in this study. MAIN OUTCOME MEASURES: Frequency, type, and strength of study design were measured. RESULTS: We evaluated 1911 articles (967 clinical; 17% comparative). There was a significant increase in multicenter clinical studies (from 12% to 27%; p < 0.0001) and mean study population (from 326 to 6775; p < 0.05). Studies using administrative data increased from 14% to 43% (p < 0.0001). Case reports decreased from 16% to 7% of all clinical studies (p < 0.001), whereas the percentage of comparative studies increased from 14% to 21% (p = 0.001). The percentage of randomized controlled trials did not increase significantly (8.5% in 2000; 10.0% in 2010; p = 0.44). The mean 10-point score for comparative studies was 6.7 for both years (p = 0.50). There was good interobserver agreement in the classification of studies (κ = 0.70) and moderate agreement in scoring comparative studies (κ = 0.47). LIMITATIONS: This descriptive study cannot fully account for the reasons behind the identified differences. CONCLUSIONS: Comparative and multicenter studies, mean study population, and the use of administrative data increased from 2000 to 2010. This suggests that increased use of administrative databases has allowed larger populations of patients from more institutions to be studied and may be more generalizable. Researchers should strive toward improving the level of evidence (see Video, Supplemental Digital Content 1, http://links.lww.com/DCR/A167).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.088 | 0.265 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.033 | 0.029 |
| Science and technology studies | 0.002 | 0.003 |
| Scholarly communication | 0.006 | 0.006 |
| Open science | 0.001 | 0.004 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".