The Role of Expert Opinion in Projecting Long-Term Survival Outcomes Beyond the Horizon of a Clinical Trial
Bibliographic record
Abstract
INTRODUCTION: Clinical trials often have short follow-ups, and long-term outcomes such as survival must be extrapolated. Current extrapolation methods often produce a wide range of survival values. To minimize uncertainty in projections, we developed a novel method that incorporates formally elicited expert opinion in a Bayesian analysis and used it to extrapolate survival in the placebo arm of DAPA-CKD, a phase 3 trial of dapagliflozin in patients with chronic kidney disease (NCT03036150). METHODS: A summary of mortality data from 13 studies that included DAPA-CKD-like populations and training on elicitation were provided to six experts. An elicitation survey was used to gather the experts' 10- and 20-year survival estimates for patients in the placebo arm of DAPA-CKD. These estimates were combined with DAPA-CKD mortality and general population mortality (GPM) data in a Bayesian analysis to extrapolate long-term survival using seven parametric distributions. Results were compared with those from standard frequentist approaches (with and without GPM data) that do not incorporate expert opinion. RESULTS: The group expert-elicited estimate for 20-year survival was 31% (lower estimate, 10%; upper estimate, 40%). In the Bayesian analysis, the 20-year extrapolated survival across the seven distributions was 14.9-39.1%, a range that was 2.4- and 1.6-fold smaller than those produced by the frequentist methods (0.0-56.9% without and 0.0-39.2% with GPM data). CONCLUSIONS: Using expert opinion in a Bayesian analysis provided a robust method for extrapolating long-term survival in the placebo arm of DAPA-CKD. The method could be applied to other populations with limited survival data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".