Clinical study reports of randomised controlled trials: an exploratory review of previously confidential industry reports
Bibliographic record
Abstract
OBJECTIVE: To explore the structure and content of a non-random sample of clinical study reports (CSRs) to guide clinicians and systematic reviewers. SEARCH STRATEGY: We searched public sources and lodged Freedom of Information requests for previously confidential CSRs primarily written by the industry for regulators. SELECTION CRITERIA: CSRs reporting sufficient information for extraction ('adequate'). PRIMARY OUTCOME MEASURES: Presence and length of essential elements of trial design and reporting and compression factor (ratio of page length for CSRs compared to its published counterpart in a scientific journal). DATA EXTRACTION: Data were extracted on standard forms and crosschecked for accuracy. RESULTS: We assembled a population of 78 CSRs (covering 90 randomised controlled trials; 144 610 pages total) dated 1991-2011 of 14 pharmaceuticals. Report synopses had a median length of 5 pages, efficacy evaluation 13.5 pages, safety evaluation 17 pages, attached tables 337 pages, trial protocol 62 pages, statistical analysis plan 15 pages and individual efficacy and safety listings had a median length of 447 and 109.5 pages, respectively. While 16 (21%) of CSRs contained completed case report forms, these were accessible to us in only one case (765 pages representing 16 individuals). Compression factors ranged between 1 and 8805. CONCLUSIONS: Clinical study reports represent a hitherto mostly hidden and untapped source of detailed and exhaustive data on each trial. They should be consulted by independent parties interested in a detailed record of a clinical trial, and should form the basic unit for evidence synthesis as their use is likely to minimise the problem of reporting bias. We cannot say whether our sample is representative and whether our conclusions are generalisable to an undefined and undefinable population of CSRs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.821 | 0.612 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.029 | 0.005 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.044 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".