Augmented two-stage estimation for treatment switching in oncology trials: Leveraging external data for improved precision
Bibliographic record
Abstract
Randomized controlled trials in oncology often allow control group participants to switch to experimental treatments, a practice that, while often ethically necessary, complicates the accurate estimation of long-term treatment effects. When switching rates are high or sample sizes are limited, commonly used methods for treatment switching adjustment (such as the rank-preserving structural failure time model, inverse probability of censoring weights, and two-stage estimation) may produce imprecise estimates. Real-world data can be used to develop an external control arm for the randomized controlled trial, although this approach ignores evidence from trial subjects who did not switch and ignores evidence from the data obtained prior to switching for those subjects who did. This article introduces "augmented two-stage estimation" (ATSE), a method that combines data from non-switching participants in a randomized controlled trial with an external dataset, forming a "hybrid non-switching arm". While aiming for more precise estimation, the augmented two-stage estimation requires strong assumptions. Namely, conditional on all the observed covariates: (1) a participant's decision to switch treatments must be independent of their post-progression survival, and (2) individuals from the randomized controlled trial and the external cohort must be exchangeable. With a simulation study, we evaluate the augmented two-stage estimation method's performance compared to two-stage estimation adjustment and an external control arm approach. Results indicate that performance is dependent on scenario characteristics, but when unconfounded external data are available, augmented two-stage estimation may result in less bias and improved precision compared to two-stage estimation and external control arm approaches. When external data are affected by unmeasured confounding, augmented two-stage estimation becomes prone to bias, but to a lesser extent compared to an external control arm approach.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.071 | 0.464 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".