MétaCan
Menu
Retour à la cohorte
Enregistrement W4414062714 · doi:10.1093/jnci/djaf183

Success in oncology phase 3 trials: a small <i>P</i> value or patient value

2025· article· en· W4414062714 sur OpenAlexaffabout
Laure-Anne Teuwen, Gregory R. Pond, Bishal Gyawali

Notice bibliographique

RevueJNCI Journal of the National Cancer Institute · 2025
Typearticle
Langueen
DomaineMathematics
ThématiqueStatistical Methods in Clinical Trials
Établissements canadiensQueen's UniversityMcMaster University
Organismes subventionnairesnon disponible
Mots-clésValue (mathematics)Phase (matter)MEDLINEPatient care

Résumé

récupéré en direct d'OpenAlex

Cancer drug development is a complex, multistep process that spans many years and costs billions of dollars—yet only a small fraction of these drugs ultimately demonstrate a clinically meaningful benefit in a phase 3 trial leading to regulatory approval.1 A phase 3 trial, usually the final barrier before regulatory approval, is also where the probability of failure is the highest with studies showing that less than half of these trials are successful. Paradoxically, phase 3 trials are also the most expensive to conduct, with a phase 3 trial typically costing $22.1 million (vs $4.5 million for a phase 1 and $11.2 million for a phase 2) and needing more patients for enrollment.2 Despite all the educational values of a “negative” phase 3 trial (negative in a purely statistical sense), a negative phase 3 cancer drug trial is a disappointment, not just financially but also because of the opportunity cost of human resources (patients) who could have been enrolled in other trials to answer other important questions. Thus, it is preferable that a negative phase 3 cancer drug trial be avoided. In the article accompanying this editorial, Lamp and Mazza3 provocatively ask if these phase 3 trials are unsuccessful because of unrealistic expectations of the researchers. They reviewed 111 phase 3 oncology randomized clinical trials published in top journals and found that a majority (57.3%) reported hazard ratios (HRs) weaker than those assumed in their respective power calculations. Lamp and Mazza3 find this problematic because if reported hazard ratios were consistently close to the point estimate, the observed hazard ratio would be weaker than assumed hazard ratio in 50% cases rather than the 57% found in their analysis. They argue that researchers tend to adopt overly optimistic assumptions about effect sizes in their power calculations, which would occur if researchers used previous phase 2 results as a basis for phase 3 hypotheses without accounting for the known bias resulting from small, positive trials, which tend to overestimate the true treatment effect.4 In turn, this would result in underpowered trials. Lamp and Mazza3 conclude that power calculations should be based on less optimistic expectations with trials testing for smaller efficacy differences between the arms and suggest using the lower bound of the confidence interval, rather than a point estimate, to inform sample size calculations for phase 3 trials. However, is lowering the hazard ratio from point estimate to lower bound of confidence interval the right approach to address this problem? We believe not. “Overoptimism” may play some minor role in phase 3 trials being negative; however, oncology’s problem has not been that we are missing out on several game-changing drugs—our problem has been quite the opposite. We are flooded with drugs that provide minimal clinical benefit and yet pass a statistical hurdle in a phase 3 trial. This problem would only be compounded with lowering the hazard ratio expectations to the lower bound of confidence interval, at the expense of increasing trial sample sizes and the costs of conducting trials. Already, the average survival gain provided by cancer drugs that are not just positive on a phase 3 but eventually approved by the US Food and Drug Administration is only 2.9 months,5 which many would agree is modest at best. To call this an “optimistic” expectation and advocate to lower the bar is definitely not in the patients’ interest. We argue that this is akin to a school claiming increased distinction rates in exams not by improving the quality of their students but by lowering the threshold of what counts as a distinction. Lamp and Mazza3 do not provide information on median survival gains but report that the average observed hazard ratio in their sample was 0.72 (vs average expected HR of 0.66). However, in their cohort, only 38% of the trials had the primary endpoint of overall survival; the rest used putative surrogate endpoints such as progression-free survival, which in most cases is a poor predictor for any patient-centric benefit such as survival or quality of life.6 There are many trials in oncology in which even large progression-free survival gains have failed to translate into survival gains. We strongly argue that current oncology trials are already overpowered and are frequently designed to detect clinically trivial differences in endpoints that do not matter. We believe that the change needed in oncology trial design is not to test for even smaller differences but to power the trial appropriately to test only for those survival gains that are clinically meaningful, similar to how quality-of-life analyses are conducted.7 This is aligned with the recommendations from the Common Sense Oncology initiative that trials should be designed to detect clinically meaningful differences as identified by validated scales such as the European Society for Medical Oncology- Magnitude of Clinical Benefit Scale scores.8 In fact, the Common Sense Oncology principles recommend that trials that fail to meet such meaningful thresholds be reported as failing to show clinically meaningful benefit (ie, negative) irrespective of statistical significance. We agree with Lamp and Mazza3 that the success rates of phase 3 trials should improve, but we don’t agree that the way to improve this is by lowering the criteria for defining success. If anything, given that most of our drugs offer only marginal benefits, oncology needs strengthening of criteria. Irrespective of the target effect size, cancer drug trials are known to chase statistical significance by overpowering the trial by enrolling a larger number of patients than indicated and testing for outcomes that are easier to achieve statistical significance rather than outcomes that are of greatest importance to patients.9 So, if lowering the expectation is not the way forward, what is? How can we avoid negative phase 3 trials? We believe that although we can certainly improve the success rates, it is not reasonable to expect that we can avoid negative phase 3 trials. By definition, a randomized trial is conducted when there is clinical equipoise between the 2 arms. Thus, by definition, we should expect our phase 3 randomized trials to yield negative results quite often. If the phase 3 randomized trials yielded positive results most of the time, the ethics of conducting these trials and randomizing patients to the control arm should be questioned. The appropriate question is, How can we reasonably improve the probability of a positive phase 3 trial? One way to do this would be to test only those drugs in a phase 3 that are more likely to demonstrate large, clinically meaningful differences. Previous studies have showed that higher phase 2 response rates were predictive of phase 3 success and eventual drug approval.10,11 However, a study of negative phase 3 trials in oncology revealed that 42% were conducted without any prior phase 2 evidence.12 These were trials that moved directly from phase 1 to phase 3, bypassing phase 2, and ended up being negative at phase 3. In fact, 28% of these negative phase 3 trials had a negative preceding phase 2. Let’s digest that for a minute. The drugs failed in a phase 2 and yet were moved to a phase 3 despite all the financial and human costs of running a phase 3 trial. The pretest probability that the phase 3 would be positive was quite low, and yet the phase 3 trial was still conducted. There are also instances where redundant and duplicative phase 3 trials have been conducted for the same or similar molecule on multiple tumor types, hoping that one of the trials would turn positive through statistical probability even if the drug was not very effective.13 These reflect the perverse incentives in the oncology marketplace. Cancer drugs, when approved, can be priced so high that the return on investment might justify running phase 3 trials hoping for a false-positive result.14 One way to address this at the policy level would be to ensure that the prices of cancer drugs are aligned with the magnitude of clinical benefit, so that a small statistically positive result is not rewarded financially the same way as a drug that offers substantial clinical benefit. In the end, it is important to not lose sight of the ultimate endpoint in cancer research. The primary endpoint for everything we do in cancer research and trials is to improve patient outcomes by prolonging the length and/or quality of the patients’ lives. It is not to improve phase 3 trial success rate or find a new biomarker or get another drug approved—these are only surrogates. The only thing that matters is patient outcomes. Laure-Anne Teuwen (Writing—original draft, Writing—review & editing), Gregory Pond(Methodology, Writing—review & editing), and Bishal Gyawali (Conceptualization, Methodology, Supervision, Writing—original draft, Writing—review & editing). No funding was used for this editorial. L.A.T. has received honoraria from AstraZeneca, unrelated to the manuscript. L.A.T. is supported by the Belgian Foundation Against Cancer. G.R.P. has received consulting fees from Traferox Technologies and Calian CRO and has a close family member who is employed by Roche Canada and who owns stock in Roche Ltd. B.G. has received consulting fees from Vivio Health, unrelated to the manuscript. B.G. acknowledges salary support from Ontario Institute for Cancer Research funded by the Government of Ontario. The opinions expressed in the editorial are his own and do not represent the views of the Government. No new data were generated or analyzed for this editorial.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,012
score de la tête « metaresearch » (Gemma)0,248
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Théorique ou conceptuel · Signal consensuel: Théorique ou conceptuel
GenreSignal candidat: Méthodes · Signal consensuel: aucune
Score de désaccord entre enseignants0,511
Score d'incertitude au seuil0,758

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0120,248
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0010,000
Bibliométrie0,0000,001
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0010,000
Intégrité de la recherche0,0000,001
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,754
Tête enseignante GPT0,669
Écart entre enseignants0,084 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Devis d'étudeThéorique ou conceptuel
Domainenon disponible
GenreMéthodes

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations2
Publié2025
Routes d'admission2
Résumé présentoui

Explorer davantage

Même revueJNCI Journal of the National Cancer InstituteMême sujetStatistical Methods in Clinical TrialsTravaux en français237 207