Efficacy-effectiveness gap with immune checkpoint inhibitors in solid cancers: A population based cohort study.
Bibliographic record
Abstract
6625 Background: Immune checkpoint inhibitors (ICIs) have been central in oncology drug development in the last decade often demonstrating large improvements in overall survival (OS) in pivotal randomized controlled trials (RCTs). Patient selection and/or care may be less stringent in the real world but ascertaining that benefits are translated satisfactorily from pivotal RCTs into the real world is crucial, as patients often base their treatment decisions on outcomes quoted to them from pivotal RCTs. Methods: After necessary ethics approvals, we collected relevant data on all consecutive patients with advanced solid cancers treated with ICIs in the province of Manitoba, Canada from inception to June 30, 2022 using pharmacy database and cancer registry. Eight most frequent ICI indications are included here. Survival probabilities were plotted using Kaplan-Meier Method, and descriptive statistics used to report relationship between variables. To put outcomes into perspective, we report real world data alongside that of pivotal RCTs used to support approvals of respective ICIs by Health Canada and the US FDA. Results: Eight ICI indications included 992 patients followed for a median of 40 months (range: 21-86 months). Patients required to meet precise provincial formulary criteria to receive ICIs, adopted mostly from the inclusion criteria of the corresponding pivotal RCTs. Median OS was shorter in the real world population not only compared to the experimental arms of the respective pivotal RCTs [by a median of 16 months (range: 3 - 38 months)] but also compared to their control arms [by a median of 3 months (range 3 to 32 months)]. This translated into an average of 51.1% (standard deviation: 20.1) compromise in median OS with ICIs in the real world compared to that observed in respective pivotal RCTs. Conclusions: Treatment with ICIs within high standards of practice in a Canadian province overall offered less than half the median OS observed in pivotal RCTs supporting their approvals. Current ICI treatment decisions by patients (and physicians) may therefore be based on highly inaccurate expectations. Our results should inspire rigorous phase IV clinical studies to identify pertinent real-world variables affecting outcomes including treatment related morbidity and mortality, with an ultimate aim to narrow the alarming efficacy-effectiveness gap. [Table: see text]
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.002 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".