Modeling Challenges in Cost-Effectiveness Analysis of First-Line Immuno-Oncology Therapies in Non-small Cell Lung Cancer: A Systematic Literature Review
Bibliographic record
Abstract
INTRODUCTION: The introduction of immuno-oncology (IO) therapies has changed the treatment landscape of non-small cell lung cancer (NSCLC). Numerous cost-effectiveness analyses (CEAs) and technology appraisals (TAs) evaluating IO therapies have been recently published. OBJECTIVE: We reviewed economic models of first-line (1L) IO therapies for previously untreated advanced or metastatic NSCLC to identify methodological challenges associated with modeling cost effectiveness from published literature and TAs and to make recommendations for future CEAs in this disease area. METHODS: A systematic literature review was conducted following Cochrane and PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines. We searched MEDLINE, Embase, EconLit (January 2009-January 2020), and select conferences (since 2016) for CEAs of 1L IO treatments in patients with recurrent or metastatic, epidermal growth factor receptor (EGFR)/anaplastic lymphoma kinase (ALK) mutation-negative NSCLC, published in English. TAs from England, Scotland, Canada, Australia, Germany, and France were also examined. Two reviewers screened the results and extracted the data. The quality of the CEAs was described using the Drummond checklist. RESULTS: In total, 46 records reporting on 38 unique models met protocol-defined criteria and were included. Five models adjusted for treatment switching or crossover in base-case analyses, and the remainder considered treatment switching or crossover to represent clinical practice and made no adjustment. Seven models used external real-world data for survival modeling or extrapolation validation. Six models that assumed long-term treatment benefit stopped at 3 or 5 years after initiation. Seven models used the observed time-on-treatment distribution from the trial, and eight used progression-free survival for treatment duration. All models compared one or more IO monotherapies or combination therapies with chemotherapy. Only one study directly compared different IO agents but did not consider the concordance issue across programmed death-ligand 1 (PD-L1) testing methods. Utilities were modeled by health state in 12 models, four applied a time-to-death approach, and ten explored both. None applied cure models. CONCLUSION: Variations in methodological challenges were seen across studies. Previous models took approaches that were followed in subsequent models, such as a 2-year stopping rule of IO duration or treatment-effect waning. Challenges such as heterogeneity in PD-L1 testing and survival extrapolation and validation using real-world data should be further considered for future models in advanced or metastatic NSCLC.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.009 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".