Inefficiencies and Patient Burdens in the Development of the Targeted Cancer Drug Sorafenib: A Systematic Review
Bibliographic record
Abstract
Failure in cancer drug development exacts heavy burdens on patients and research systems. To investigate inefficiencies and burdens in targeted drug development in cancer, we conducted a systematic review of all prelicensure trials for the anticancer drug, sorafenib (Bayer/Onyx Pharmaceuticals). We searched Embase and MEDLINE databases on October 14, 2014, for prelicensure clinical trials testing sorafenib against cancers. We measured risk by serious adverse event rates, benefit by objective response rates and survival, and trial success by prespecified primary endpoint attainment with acceptable toxicity. The first two clinically useful applications of sorafenib were discovered in the first 2 efficacy trials, after five drug-related deaths (4.6% of 108 total) and 93 total patient-years of involvement (2.4% of 3,928 total). Thereafter, sorafenib was tested in 26 indications and 67 drug combinations, leading to one additional licensure. Drug developers tested 5 indications in over 5 trials each, comprising 56 drug-related deaths (51.8% of 108 total) and 1,155 patient-years (29.4% of 3,928 total) of burden in unsuccessful attempts to discover utility against these malignancies. Overall, 32 Phase II trials (26% of Phase II activity) were duplicative, lacked appropriate follow-up, or were uninformative because of accrual failure, constituting 1,773 patients (15.6% of 11,355 total) participating in prelicensure sorafenib trials. The clinical utility of sorafenib was established early in development, with low burden on patients and resources. However, these early successes were followed by rapid and exhaustive testing against various malignancies and combination regimens, leading to excess patient burden. Our evaluation of sorafenib development suggests many opportunities for reducing costs and unnecessary patient burden in cancer drug development.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".