Delivery of meaningful cancer care: Evaluating benefit and cost of cancer therapies using ASCO and ESMO frameworks.
Bibliographic record
Abstract
6609 Background: ASCO and ESMO have developed frameworks to evaluate the benefit of cancer therapies. Here, we apply the frameworks to a cohort of contemporary randomized controlled trials (RCTs) to explore agreement and to evaluate the relationship between treatment benefit and cost. Methods: Characteristic and outcome data from RCTs evaluating systemic therapies in non-small cell lung cancer (NSCLC), breast cancer, colorectal cancer (CRC), and pancreatic cancer published and cited in PubMed between 2011-2015 were abstracted. Trial endpoints were evaluated using ASCO and ESMO frameworks. Cohen’s kappa statistic was calculated to determine agreement between the two frameworks, using the median ASCO score as a benefit threshold. Differences in monthly drug cost between RCT experimental and control arms were derived from 2016 average wholesale prices. Analyses included Pearson chi-square tests, Fisher’s Exact tests, independent samples t-tests, and Pearson correlation to assess the association between continuous variables. Results: Fifty percent (136/271) of published RCTs favoured the experimental arm; scoring rubrics were applicable to 109 RCTs (39% NSCLC, 33% breast, 23% CRC, 5% pancreas). ASCO scores ranged from 2 to 72; median score was 25. Thirty seven percent (40/109) of RCTs met benefit thresholds using the ESMO framework. Agreement between frameworks was fair at best (κ = 0.28, p = 0.002). When stratified by treatment intent (19 curative, 90 palliative RCTs), agreement remained poor (κ = 0.23, p = 0.115; κ = 0.34, p < 0.001). Major differences leading to limited agreement includes the relative weights each framework places on HR, endpoints, and toxicity/QOL analysis. Smaller RCT sample size was the only trial characteristic associated with higher ASCO scores (p = 0.015). Among the 100 RCTs for whom drug costing data were available, there was no association between ASCO benefit score and monthly drug costs (r = -0.12, p = 0.22); those meeting ESMO thresholds had a lower mean drug cost than those who did not (p = 0.046). Conclusions: There is only fair correlation between ASCO and ESMO clinical benefit frameworks. Drug costs are not associated with ESMO/ASCO measures of magnitude of clinical benefit.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.243 | 0.408 |
| Meta-epidemiology (narrow) | 0.003 | 0.001 |
| Meta-epidemiology (broad) | 0.006 | 0.017 |
| Bibliometrics | 0.029 | 0.017 |
| Science and technology studies | 0.001 | 0.004 |
| Scholarly communication | 0.007 | 0.005 |
| Open science | 0.003 | 0.010 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".