Avoiding the hazards of misinterpreting treatment effects
Bibliographic record
Abstract
Time-to-event end points are the most frequent outcome measures in oncology phase III trials. Since these trials inform clinical practice, a clear understanding of the measures used to describe the effect of treatment on such outcomes is important for oncologists, who need to communicate appropriately with patients and their families [1.Saad E.D. Zalcberg J.R. Peron J. et al.Understanding and communicating measures of treatment effect on survival: can we do better?.J Natl Cancer Inst. 2018; 110: 232-240Crossref PubMed Scopus (35) Google Scholar]. The effect of treatment can be expressed in absolute and in relative terms, depending on individual preference, the frequency of events, and accepted standards in a given field. It has been known for decades that physicians have differing perceptions about the effect of treatment depending on whether absolute or relative measures are presented [2.Naylor C.D. Chen E. Strauss B. Measured enthusiasm: does the method of reporting trial results alter perceptions of therapeutic effectiveness?.Ann Intern Med. 1992; 117: 916-921Crossref PubMed Scopus (336) Google Scholar, 3.Bobbio M. Demichelis B. Giustetto G. Completeness of reporting trial results: effect on physicians' willingness to prescribe.Lancet. 1994; 343: 1209-1211Abstract PubMed Scopus (177) Google Scholar]. The way trial results are communicated by clinicians may influence the acceptance of different treatments by people with cancer [4.Chao C. Studts J.L. Abell T. et al.Adjuvant chemotherapy for breast cancer: how presentation of recurrence risk influences decision-making.J Clin Oncol. 2003; 21: 4299-4305Crossref PubMed Scopus (76) Google Scholar]. In this issue of Annals of Oncology, Weir et al. describe the results of a survey conducted to assess how reporting the hazard ratio (HR) or the difference in restricted mean survival times (RMSTD) might influence the interpretation of the treatment effect by physicians [5.Weir I.R. Marshall G.D. Schneider J.I. et al.Interpretation of time-to-event outcomes in randomized trials: an online randomized experiment.Ann Oncol. 2019; 30: 96-102Abstract Full Text Full Text PDF PubMed Scopus (25) Google Scholar]. These authors have used an ingenious study design, stratifying, and randomizing three groups of participants to receive one of the three different versions of an abstract reporting treatment effects using the HR, the RMSTD, or both. Each abstract was selected from 15 publications describing the outcomes of real phase III trials, but it was modified to de-identify its origin and to present the treatment effect in one of the three ways proposed by the authors. Participants were then asked to rate the treatment benefit on a scale, with lower values indicating less benefit. The thoughtfulness of this study is exemplified by computation of the HR for each trial based on reconstructed individual-patient data, to be consistent with computation of restricted means. As a secondary objective, Weir et al. sought to assess the extent to which the HR is interpreted correctly by clinical trialists, fellows, and residents. The key finding of the study was that participants judged the treatment effect to be more pronounced when the HR was reported, either alone or with the RMSTD, than when the RMSTD was reported alone. Counterintuitively, differences in interpretation were nominally smaller among fellows and residents than among clinical trialists, and nominally larger among those with a degree in epidemiology or biostatistics. Nearly half of the participants misinterpreted the HR, with ∼40% equating it to relative risk. A caveat for these findings is representativeness, because the participation rate was only 7%, despite the laudable strategy of rewarding acceptance to participate by donation to a charitable fund. There are two subtle issues that may have influenced the results. The first is familiarity of the medical community with the HR, and lack of familiarity with restricted means. Even though the HR is often misunderstood, familiarity with this measure may have led some participants to rate the treatment benefit of the trials ‘away from the null hypothesis’ more often than was the case for participants assessing the RMSTD. Secondly, the perception of larger benefit seen when the HR was presented may have occurred because relative measures (e.g. the HR) result in larger numerical magnitudes of treatment effect than their absolute counterparts (e.g. the RMSTD), thus conveying the impression of a larger benefit [2.Naylor C.D. Chen E. Strauss B. Measured enthusiasm: does the method of reporting trial results alter perceptions of therapeutic effectiveness?.Ann Intern Med. 1992; 117: 916-921Crossref PubMed Scopus (336) Google Scholar, 3.Bobbio M. Demichelis B. Giustetto G. Completeness of reporting trial results: effect on physicians' willingness to prescribe.Lancet. 1994; 343: 1209-1211Abstract PubMed Scopus (177) Google Scholar]. It would be interesting to contrast the perception of benefit conveyed by the HR and its relative counterpart, the ratio of restricted mean survival times (or, conversely, by the RMSTD and its absolute counterpart, the difference in medians). Interestingly, the same group has previously shown that the HR provides a larger numerical magnitude of treatment effect than does the ratio of restricted mean survival times [6.Trinquart L. Jacot J. Conner S.C. Porcher R. Comparison of treatment effects measured by the hazard ratio and by the ratio of restricted mean survival times in oncology randomized controlled trials.J Clin Oncol. 2016; 34: 1813-1819Crossref PubMed Scopus (155) Google Scholar]. Weir et al. are not alone in revealing misunderstanding of statistical concepts in medical circles, something that goes beyond measures of treatment effect [7.Wasserstein R.L. Lazar NA. The ASA's statement on p-values: context, process, and purpose.Am Stat. 2016; 70: 129-133Crossref Scopus (3258) Google Scholar, 8.Greenland S. Senn S.J. Rothman K.J. et al.Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations.Eur J Epidemiol. 2016; 31: 337-350Crossref PubMed Scopus (1324) Google Scholar]. The American Statistical Association has issued a statement emphasizing the misuse of P-values, which should not be used alone to make clinical decisions (Box 1) [7.Wasserstein R.L. Lazar NA. The ASA's statement on p-values: context, process, and purpose.Am Stat. 2016; 70: 129-133Crossref Scopus (3258) Google Scholar]. Nevertheless, major journals and regulatory agencies often do exactly that in their requirements for the publication of trials and for the registration of new treatments. A more holistic evaluation of outcomes in clinical trials is needed that emphasizes the clinical value of new treatments, as in efforts by the European Society of Medical Oncology and the American Society of Clinical Oncology [9.Cherny N.I. Sullivan R. Dafni U. et al.A standardised, generic, validated approach to stratify the magnitude of clinical benefit that can be anticipated from anti-cancer therapies: the European Society for Medical Oncology Magnitude of Clinical Benefit Scale (ESMO-MCBS).Ann Oncol. 2015; 26: 1547-1573Abstract Full Text Full Text PDF PubMed Scopus (532) Google Scholar, 10.Schnipper L.E. Davidson N.E. Wollins D.S. et al.American Society of Clinical Oncology Statement: a conceptual framework to assess the value of cancer treatment options.J Clin Oncol. 2015; 33: 2563-2577Crossref PubMed Scopus (679) Google Scholar]. Of major concern, the understanding of outcomes has penetrated minimally among those designing and reporting potentially practice-changing clinical trials, the referees and editors of journals that accept these reports, and those that recommend approval of new treatments at the United States Food and Drug Administration and the European Medicines Agency. Almost all phase III trials are planned to show or rule out an HR that favours an experimental treatment compared with a standard (typically, HR < 0.8 or <0.75), with a P-value for this result of <0.05. The difference between the arms in median survival (or other time-to-event primary end point) is usually quoted. Survival curves are tested rarely for the assumption of proportional hazards, and values of HR are even quoted for survival curves that cross [11.Mok T.S. Wu Y.L. Thongprasert S. et al.Gefitinib or carboplatin-paclitaxel in pulmonary adenocarcinoma.N Engl J Med. 2009; 361: 947-957Crossref PubMed Scopus (7030) Google Scholar]. The HR does not represent appropriately results of trials in which a small proportion of patients have prolonged survival so that the survival curves only separate after some time (thus not satisfying proportional hazards, as in some immunotherapy trials) [12.Liang F. Zhang S. Wang Q. Li W. Treatment effects measured by restricted mean survival time in trials of immune checkpoint inhibitors for cancer.Ann Oncol. 2018; 29: 1320-1324Abstract Full Text Full Text PDF PubMed Scopus (30) Google Scholar]; although there may be no difference in median survival or RMSTD for such trials, there is benefit for a minority of participants. For people with advanced incurable cancer, the survival curves will always come together after sufficient follow-up (banana curves)—again violating the assumption of proportional hazards, regardless of divergence of curves at earlier times. We should assess treatment benefit in these and other situations with non-proportional hazards using traditional and novel measures [1.Saad E.D. Zalcberg J.R. Peron J. et al.Understanding and communicating measures of treatment effect on survival: can we do better?.J Natl Cancer Inst. 2018; 110: 232-240Crossref PubMed Scopus (35) Google Scholar, 6.Trinquart L. Jacot J. Conner S.C. Porcher R. Comparison of treatment effects measured by the hazard ratio and by the ratio of restricted mean survival times in oncology randomized controlled trials.J Clin Oncol. 2016; 34: 1813-1819Crossref PubMed Scopus (155) Google Scholar, 13.Peron J. Roy P. Ozenne B. et al.The net chance of a longer survival as a patient-oriented measure of treatment benefit in randomized clinical trials.JAMA Oncol. 2016; 2: 901-905Crossref PubMed Scopus (37) Google Scholar, 14.Uno H. Claggett B. Tian L. et al.Moving beyond the hazard ratio in quantifying the between-group difference in survival analysis.J Clin Oncol. 2014; 32: 2380-2385Crossref PubMed Scopus (393) Google Scholar].Box 1Six principles in the statement by the American Statistical Association indicating the limitations of using P-values to describe outcomes of medical and scientific research [7.Wasserstein R.L. Lazar NA. The ASA's statement on p-values: context, process, and purpose.Am Stat. 2016; 70: 129-133Crossref Scopus (3258) Google Scholar].•P-values can indicate how incompatible the data are with a specified statistical model.•P-values do not measure the probability that the studied hypothesis is true, or the probability that the data were produced by random chance alone.•Scientific conclusions and business or policy decisions should not be based only on whether a P-value passes a specific threshold.•Proper inference requires full reporting and transparency.•A P-value, or statistical significance, does not measure the size of an effect or the importance of a result.•By itself, a P-value does not provide a good measure of evidence regarding a model or hypothesis. •P-values can indicate how incompatible the data are with a specified statistical model.•P-values do not measure the probability that the studied hypothesis is true, or the probability that the data were produced by random chance alone.•Scientific conclusions and business or policy decisions should not be based only on whether a P-value passes a specific threshold.•Proper inference requires full reporting and transparency.•A P-value, or statistical significance, does not measure the size of an effect or the importance of a result.•By itself, a P-value does not provide a good measure of evidence regarding a model or hypothesis. The study by Weir et al. adds to mounting evidence that we need to educate ourselves and our colleagues towards better understanding measures of treatment effect in oncology. In agreement with Weir et al., we encourage authors of clinical trials to report different measures of treatment effect [1.Saad E.D. Zalcberg J.R. Peron J. et al.Understanding and communicating measures of treatment effect on survival: can we do better?.J Natl Cancer Inst. 2018; 110: 232-240Crossref PubMed Scopus (35) Google Scholar, 15.A'Hern RP. Restricted mean survival time: an obligatory end point for time-to-event analysis in cancer trials?.J Clin Oncol. 2016; 34: 3474-3476Crossref PubMed Scopus (87) Google Scholar, 16.Pocock S.J. Ariti C.A. Collier T.J. Wang D. The win ratio: a new approach to the analysis of composite endpoints in clinical trials based on clinical priorities.Eur Heart J. 2012; 33: 176-182Crossref PubMed Scopus (256) Google Scholar, 17.Royston P. Parmar MK. Restricted mean survival time: an alternative to the hazard ratio for the design and analysis of randomized trials with a time-to-event outcome.BMC Med Res Methodol. 2013; 13: 152.Crossref PubMed Scopus (441) Google Scholar]. This strategy would allow better understanding of the nuances of the treatment effect in different settings and more appropriate communication with our patients and their families. We thank Dr Ronald L. Wasserstein for the permission to use material from the American Statistical Association. None declared.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".