Late, never or non‐existent: the inaccessibility of preclinical evidence for new drugs
Bibliographic record
Abstract
BACKGROUND AND PURPOSE: Animal studies establish much of the evidence used to support clinical development of new drugs. Recent studies suggest that many preclinical investigations are withheld from publication, leading to exaggerated estimates of clinical utility. We sought to estimate the volume and properties of all published animal efficacy studies for a cohort of novel drugs. EXPERIMENTAL APPROACH: We searched biomedical databases to identify 47 novel drugs whose first trials were reported between 2000 and 2003, inclusive. Next, we searched for all published animal studies testing the same drug, regardless of publication date. We then extracted items from titles and abstracts of eligible studies. KEY RESULTS: We identified 2462 efficacy studies, representing an average of 52 studies per drug. No published efficacy studies were available for three drugs in our sample. The volume of efficacy studies was related to how far the drug had progressed in clinical development (Spearman's correlation coefficient = 0.66, P < 0.0001). Most (87%) accessible animal efficacy studies were reported after publication of the first trial, and for 17% of the drugs in our sample, no efficacy studies were published before the first trial report. Disease indications used in trials often did not match those modelled in efficacy studies; for 35% of indications tested in trials, we were unable to identify any published efficacy studies in models of the same indication. CONCLUSIONS AND IMPLICATIONS: The volume of published efficacy studies is large, although numerous gaps reflect non-publication, publication delay or non-performance of efficacy studies supporting trials.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".