Re-aligning the ASCO and ESMO clinical benefit frameworks for modern cancer therapies
Bibliographic record
Abstract
We read with interest the report by Cherny et al. [1.Cherny N.I. Dafni U. Bogaerts J. et al.ESMO-magnitude of Clinical Benefit Scale version 1.1.Ann Oncol. 2017; 28: 2340-2366Abstract Full Text Full Text PDF PubMed Scopus (330) Google Scholar] describing the updated European Society for Medical Oncology Magnitude of Clinical Benefit Scale (ESMO-MCBS v1.1). We have previously noted only fair correlation between the American Society of Clinical Oncology Value Framework-Value Framework (ASCO-VF) and the original ESMO-MCBS (v1.0) outputs [2.Del Paggio J.C. Sullivan R. Schrag D. et al.Delivery of meaningful cancer care: a retrospective cohort study assessing cost and benefit with the ASCO and ESMO frameworks.Lancet Oncol. 2017; 18: 887-894Abstract Full Text Full Text PDF PubMed Scopus (70) Google Scholar]. With the updated framework we explored (i) the extent to which revisions to the updated ESMO-MCBS (v1.1) modified ‘clinically meaningful’ therapies in our previously published cohort and (ii) if there was now a stronger correlation between the ESMO-MCBS and ASCO-VF [1.Cherny N.I. Dafni U. Bogaerts J. et al.ESMO-magnitude of Clinical Benefit Scale version 1.1.Ann Oncol. 2017; 28: 2340-2366Abstract Full Text Full Text PDF PubMed Scopus (330) Google Scholar, 3.Schnipper L.E. Davidson N.E. Wollins D.S. et al.Updating the American Society of Clinical Oncology value framework: revisions and reflections in response to comments received.J Clin Oncol. 2016; 34: 2925-2934Crossref PubMed Scopus (437) Google Scholar]. We used our previously described cohort of all randomized controlled trials (RCTs) published 2011–2015 of systemic therapies for non-small-cell lung cancer, breast cancer, colorectal cancer, and pancreatic cancer [2.Del Paggio J.C. Sullivan R. Schrag D. et al.Delivery of meaningful cancer care: a retrospective cohort study assessing cost and benefit with the ASCO and ESMO frameworks.Lancet Oncol. 2017; 18: 887-894Abstract Full Text Full Text PDF PubMed Scopus (70) Google Scholar, 4.Del Paggio J.C. Azariah B. Sullivan R. et al.Do contemporary randomized controlled trials meet ESMO thresholds for meaningful clinical benefit?.Ann Oncol. 2016; 28: 157-162Abstract Full Text Full Text PDF Scopus (60) Google Scholar]. Trial end points, quality of life, and toxicity data were tabulated, and ASCO-VF and ESMO-MCBS (v1.0 and v1.1) were derived discretely for each trial. For this analysis, we sought additional quality-of-life data if trials mentioned that such data was published elsewhere—different from our previous publication, but important since these scores/grades were considered more ‘complete’ [5.Del Paggio J.C. Addressing the quality of the ESMO-MCBS.Ann Oncol. 2017; 28: 1406Abstract Full Text Full Text PDF PubMed Scopus (4) Google Scholar]. Agreement was determined using Cohen’s κ statistic and was calculated to the median ASCO score (as an arbitrary threshold of benefit) and ESMO-MCBS threshold (grades A, B, 4, and 5, explicit in the framework). After re-analysis of our original cohort using ESMO-MCBS v1.0, 8% of trials (9/109) had differing grades/scores from our previous publication [2.Del Paggio J.C. Sullivan R. Schrag D. et al.Delivery of meaningful cancer care: a retrospective cohort study assessing cost and benefit with the ASCO and ESMO frameworks.Lancet Oncol. 2017; 18: 887-894Abstract Full Text Full Text PDF PubMed Scopus (70) Google Scholar]. In comparison to ESMO-MCBS v1.0 grades, ESMO-MCBS v1.1 grades differed in 7% of trials (8/109), as outlined in Table 1: three upgrades and five downgrades, four of which resulted in a change in threshold of benefit. Median ASCO score remained 25, which was considered our arbitrary ‘threshold’ of benefit for ASCO-VF.Table 1Published randomized controlled trials of non-small-cell lung cancer (NSCLC), breast cancer, colorectal cancer (CRC), and pancreatic cancer with discordant grades between the original European Society for Medical Oncology Magnitude of Clinical Benefit Scale (ESMO-MCBS) (v1.0) and the updated ESMO-MCBS (v1.1)Trial PMIDExperimental therapy and disease siteEnd point evaluatedAbsolute end point gain and HR (with 95% CI)Toxicity analysisaToxicity analysis focused on grade 3 or greater, nonlaboratory toxicities.QOL analysisESMO-MCBS v1.0 gradeESMO-MCBS v1.1 gradeReason for grade change in v1.122674612Gefitinib for NSCLCPFS5 months; 0.54 (0.37–0.79)BalancedBalanced symptom improvement34New ‘tail of the curve’ upgrade for durable survival23020162Trastuzumab emtansine for breastOS5.8 months; 0.68 (0.55–0.85)Higher in controlImproved53New grading form for control arm OS > 24 months23078958Erlotinib for NSCLCOS in pre-specified subgroup2.1 months; HR 0.76 (0.63–0.92)Higher in experimentalImproved34Absolute survival gains between grades modified23890767Viscum album for pancreaticOS2.1 months; 0.43 (0.28–0.65)BalancedFewer disease-related symptoms34Absolute survival gains between grades modified24131140Nab-paclitaxel plus gemcitabine for pancreaticOS1.8 months and 5% at 2 years; 0.72 (0.62–0.83)Higher in experimentalNot measured32Percentage survival gain no longer applicable to lower grades25088940FOLFIRI and cetuximab for CRCOS3.7 months; 0.77 (0.62–0.96)Higher in experimentalNot mentioned31New grading form for control arm OS > 24 months25877855Ramucirumab and FOLFIRI for CRCOS1.6 months; 0.844 (0.73–0.976)BalancedNo improvement21Boolean operators between HR and absolute survival gains modified26338525FOLFOXIRI and bevacizumab for CRCOS4 months; 0.80 (0.65–0.98)Higher in experimentalNot mentioned32New grading form for control arm OS > 24 monthsCI, confidence interval; FOLFIRI, folinic acid, 5-fluorouracil, and irinotecan; FOLFOXIRI, folinic acid, 5-fluorouracil, oxaliplatin, and irinotecan; HR, hazard ratio; NSCLC, non-small-cell lung cancer; OS, overall survival; PFS, progression-free survival; PMID, PubMed reference number; QoL, quality of life.a Toxicity analysis focused on grade 3 or greater, nonlaboratory toxicities. Open table in a new tab CI, confidence interval; FOLFIRI, folinic acid, 5-fluorouracil, and irinotecan; FOLFOXIRI, folinic acid, 5-fluorouracil, oxaliplatin, and irinotecan; HR, hazard ratio; NSCLC, non-small-cell lung cancer; OS, overall survival; PFS, progression-free survival; PMID, PubMed reference number; QoL, quality of life. Agreement between ESMO-MCBS v1.0 and ASCO-VF frameworks was fair (κ = 0.398); within the curative subset of trials, the κ score was weaker (0.231), with stronger agreement in the palliative subset (κ = 0.428). These results are consistent with our previously published results [2.Del Paggio J.C. Sullivan R. Schrag D. et al.Delivery of meaningful cancer care: a retrospective cohort study assessing cost and benefit with the ASCO and ESMO frameworks.Lancet Oncol. 2017; 18: 887-894Abstract Full Text Full Text PDF PubMed Scopus (70) Google Scholar]. For the updated ESMO-MCBS v1.1 and ASCO-VF, the same correlation was observed (κ = 0.397 for the full cohort, κ = 0.231 for the curative subset, and κ = 0.427 for the palliative subset). The stability of the grades between the two iterations of the framework (i.e. minimal 7% change) is consistent with that of the ESMO-MCBS authorship group, who found a total cohort grade change of 10% (12/118 grades) [1.Cherny N.I. Dafni U. Bogaerts J. et al.ESMO-magnitude of Clinical Benefit Scale version 1.1.Ann Oncol. 2017; 28: 2340-2366Abstract Full Text Full Text PDF PubMed Scopus (330) Google Scholar]. Despite these changes, agreement between the updated ESMO-MCBS and ASCO-VF in our cohort remains fair. There is a weaker agreement between ESMO-MCBS and ASCO-VF for curative scoring compared with palliative scoring. Given these findings, modifying the way in which each framework credits ‘cure’ from a therapy may be a start towards harmonization; that is, if harmonization is even a necessary goal. The algorithmic differences between value frameworks will inherently provide users with a unique ‘vantage point’ on the value of any anticancer therapy. Acknowledging that these frameworks differ at their core is what is arguably most important for clinicians applying these framework outputs to particular clinical settings: patient decision-making, trial design, guideline development, or public policy. Value frameworks—like value, itself—are relative. None declared.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".