Impact of Limited Sample Size and Follow-up on Partitioned Survival and Multistate Modeling-Based Health Economic Models: A Simulation Study
Bibliographic record
Abstract
BackgroundEconomic models often require extrapolation of clinical time-to-event data for multiple events. Two modeling approaches in oncology that incorporate time dependency include partitioned survival models (PSM) and semi-Markov decision models estimated using multistate modeling (MSM). The objective of this simulation study was to assess the performance of PSM and MSM across datasets with varying sample size and degrees of censoring.MethodsWe generated disease trajectories of progression and death for multiple hypothetical populations with advanced cancers. These populations served as the sampling pool for simulated trial cohorts with multiple sample sizes and various levels of follow-up. We estimated MSM and PSM by fitting survival models to these simulated datasets with different approaches to incorporating general population mortality (GPM) and selected best-fitting models using statistical criteria. Mean survival was compared with "true" population values to assess error.ResultsWith near complete follow-up, both PSMs and MSMs accurately estimated mean population survival, while smaller samples and shorter follow-up times were associated with a larger error across approaches and clinical scenarios, especially for more distant clinical endpoints. MSMs were slightly more often not estimable when informed by studies with small sample sizes or short follow-up, due to low numbers at risk for the downstream transition. However, when estimable, the MSM models more commonly produced a smaller error in mean survival than the PSMs did.ConclusionsCaution should be taken with all modeling approaches when the underlying data are very limited, particularly PSMs, due to the large errors produced. When estimable and for selections based on statistical criteria, MSMs performed similar to or better than PSMs in estimating mean survival with limited data.HighlightsCaution should be taken with all modeling approaches when underlying data are very limited.Partitioned survival models (PSMs) can lead to significant errors, particularly with limited follow-up. Incorporating general population mortality (GPM) via internal additive hazards improved estimates of mean survival, but the effects were modest.When estimable, decision models based on multistate modeling (MSM) produced similar or smaller error in mean survival compared with PSM, but small samples or limited deaths after progression produce additional challenges for fitting MSMs; more research is needed to improve estimation of MSMs and similar state transition-based modeling methods with limited data.Future studies are needed to assess the applicability of these findings to comparative analyses estimating incremental survival benefits.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.013 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".