Overestimation of survival in individuals with cancer treated in the real world with new cancer drugs: A population-based cohort study.
Bibliographic record
Abstract
e23076 Background: Oncologists and individuals with cancer often rely on outcomes from pivotal randomized controlled trials (RCTs) for real-world treatment decisions. However, stringent inclusions and standardized protocols limit their applicability. This study quantifies overall survival (OS) disparities for novel cancer drugs between real-world and pivotal RCTs. Methods: Ethics approvals were secured, and data on frequently used targeted cancer drugs for solid cancer treatment in all individuals with cancer in Manitoba, Canada from inception to June 30, 2023, collected using pharmacy database and the Manitoba Cancer Registry. Evidence consistently indicates similar cancer outcomes in Manitoba to most high-resource countries. This analysis includes six common immunotherapies and monoclonal antibodies. Kaplan-Meier Method was used to plot survival probabilities, and descriptive statistics to elucidate variable differences. Comparison between survival outcomes in pivotal RCTs and real-world scenarios was expressed as "Overestimation Quotient or OQ" (OS in RCT over OS in the real world for same indication). Results: A total of 941 individuals diagnosed with cancer were evaluated across eight indications. The observed median OS in the real world for included indications was 11.9 months, ranging from 6.6 to 37.6 months— a duration significantly shorter than the median OS reported in pivotal Randomized Controlled Trials (RCTs) (31.35 months, range 9.2 to 72.1) (P < 0.001). This translated to a median difference of 15.6 months (range 2.6 to 59.9) per indication and an "Overestimation Quotient" ranging from 1.5 to 5.9. Notably, the observed OS was even inferior in the real-world setting compared to the control group used in pivotal RCTs in half of the indications (Table 1). Conclusions: In routine practice, new cancer drugs offer starkly inferior survival compared to that reported in pivotal RCTs used for approval of those drugs, with the difference in survival ranging from a few months to several years, translating up to six-fold difference. Our results highlight potential for alarming misinformation influencing treatment decisions in clinics. We trust that oncologists will consider our findings during treatment discussions in clinics. [Table: see text]
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.011 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".