The randomised controlled trial in medical research: using bibliometric methods to identify core journals
Bibliographic record
Abstract
A review of: Tsay, Migh-yueh, and Yen-hsu Yang. “Bibliometric Analysis of the Literature of Randomized Controlled Trials.” Journal of the Medical Library Association 93.4 (October 2005): 450-58. Objective – To explore the characteristics and distribution of randomized controlled trials (RCTs) in the medical literature. The study aims to identify the growth patterns of the RCT, key subject matter, country and language of publication, and determine a list of core journals which contain a substantial proportion of the RCT literature. Design – Retrospective analysis of RCTs. Setting – Medical journal literature. Subjects – A total of 160,213 articles published between 1965-2001. Detailed analysis of a subset numbering 114,850 articles published from 1990-2001. Methods – The study seeks to identify all RCTs in MEDLINE from 1965-2001, and examines the growth rate of the RCT. The authors then do a more detailed analysis on a subset of data from 1990-2001, using Access database and Excel spreadsheet software, and PERL programming language. The references were analyzed by five fields within MEDLINE; publication type, source, language, country of publication, and descriptor (subject index). Main results – An exponential growth rate for the RCT is demonstrated, suggesting that in the medical literature development has not yet matured and that research using this method continues to grow. A growth rate for the RCT of 11.2% per annum is identified. The most common form of publication is the journal article, making up approximately 98% of the RCT literature. Approximately 75% of the RCTs are multicentre trials indicating that this is the design of choice adopted by researchers. The United States proves to be the greatest source of RCT literature, with 39.9% of journals and 50.6% of articles originating there. After the USA, the most productive countries are England (15.8% of journals and 21.7% articles) and Germany (6.5% journals and 6.1% articles). As might be expected, English is the predominant language providing 92.9% of the total publications. Of the remaining 7%, German is the most common language accounting for 2.2%. The top three areas being researched are: 1. Drug therapy for hypertension - 2291 citations 2. Anticancer drug combinations - 2140 citations 3. Drug therapy and asthma - 1397 citations Bradford’s law of scattering is successfully applied, identifying four zones of journals which each publish approximately 26,000 articles. Conclusion – The results indicate that bibliometric methods can be applied to the medical literature, and highlight those disciplines in which RCTs more often occur. A core list of 42 journal titles is presented, providing busy practitioners with invaluable guidance as to which journals are most likely to publish the greater number of RCTs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.269 | 0.680 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.011 | 0.007 |
| Bibliometrics | 0.219 | 0.249 |
| Science and technology studies | 0.003 | 0.003 |
| Scholarly communication | 0.017 | 0.015 |
| Open science | 0.003 | 0.007 |
| Research integrity | 0.003 | 0.003 |
| Insufficient payload (model declined to judge) | 0.006 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".