Challenges with Estimating Long-Term Overall Survival in Extensive Stage Small-Cell Lung Cancer: A Validation-Based Case Study
Bibliographic record
Abstract
Sukhvinder Johal,1 Lance Brannman,2 Victor Genestier,3 Hélène Cawston4 1Oncology Market Access and Pricing, AstraZeneca, Cambridge, UK; 2Oncology Market Access and Pricing, AstraZeneca, Gaithersburg, MD, USA; 3Health Economic and Outcomes Research, Amaris Consulting, Toronto, Ontario, Canada; 4Health Economic Outcomes Research, Amaris Consulting, Paris, FranceCorrespondence: Sukhvinder Johal, Oncology Market Access and Pricing, AstraZeneca, Cambridge, UK, Tel +44 7384 905033, Email sukhvinder.johal@astrazeneca.comObjective: The study aimed to explore methods and highlight the challenges of extrapolating the overall survival (OS) of immunotherapy-based treatment in first-line extensive stage small-cell lung cancer (ES-SCLC).Methods: Standard parametric survival models, spline models, landmark models, mixture and non-mixture cure models, and Markov models were fitted to 2-year data of the CASPIAN Phase 3 randomised trial of PD-L1 inhibitor durvalumab added to platinum-based chemotherapy (NCT03043872). Extrapolations were compared with updated 3-year data from the same trial and the plausibility of long-term estimates assessed.Results: All models used provided a reasonable fit to the observed Kaplan–Meier (K-M) survival data. The model which provided the best fit to the updated CASPIAN data was the mixture cure model. In contrast, the landmark analysis provided the least accurate fit to model survival. Estimated mean OS differed substantially across models and ranged from (in years) 1.41 (landmark model) to 4.81 (mixture cure model) for durvalumab plus etoposide and platinum and from 1.01 (landmark model) to 2.00 (mixture cure model) for etoposide and platinum.Conclusion: While most models may provide a good fit to K-M data, it is crucial to assess beyond the statistical goodness-of-fit and consider the clinical plausibility of the long-term predictions. The more complex cure models demonstrated the best predictive ability at 3 years, potentially providing a better representation of the underlying method of action of immunotherapy; however, consideration of the models’ clinical plausibility and cure assumptions need further research and validation. Our findings underscore the significance of adopting a clinical perspective when selecting the most appropriate approach to model long-term survival, particularly when considering the use of more complex models.Keywords: survival analysis, parametric extrapolation, spline model, cure models, landmark model, extensive stage small-cell lung cancer
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".