Surrogate endpoints for overall survival in chemotherapy and radiotherapy trials in operable and locally advanced lung cancer: a re-analysis of meta-analyses of individual patients' data
Bibliographic record
Abstract
BACKGROUND: The gold standard endpoint in clinical trials of chemotherapy and radiotherapy for lung cancer is overall survival. Although reliable and simple to measure, this endpoint takes years to observe. Surrogate endpoints that would enable earlier assessments of treatment effects would be useful. We assessed the correlations between potential surrogate endpoints and overall survival at individual and trial levels. METHODS: We analysed individual patients' data from 15,071 patients involved in 60 randomised clinical trials that were assessed in six meta-analyses. Two meta-analyses were of adjuvant chemotherapy in non-small-cell lung cancer, three were of sequential or concurrent chemotherapy, and one was of modified radiotherapy in locally advanced lung cancer. We investigated disease-free survival (DFS) or progression-free survival (PFS), defined as the time from randomisation to local or distant relapse or death, and locoregional control, defined as the time to the first local event, as potential surrogate endpoints. At the individual level we calculated the squared correlations between distributions of these three endpoints and overall survival, and at the trial level we calculated the squared correlation between treatment effects for endpoints. FINDINGS: In trials of adjuvant chemotherapy, correlations between DFS and overall survival were very good at the individual level (ρ(2)=0.83, 95% CI 0.83-0.83 in trials without radiotherapy, and 0.87, 0.87-0.87 in trials with radiotherapy) and excellent at trial level (R(2)=0.92, 95% CI 0.88-0.95 in trials without radiotherapy and 0.99, 0.98-1.00 in trials with radiotherapy). In studies of locally advanced disease, correlations between PFS and overall survival were very good at the individual level (ρ(2) range 0.77-0.85, dependent on the regimen being assessed) and trial level (R(2) range 0.89-0.97). In studies with data on locoregional control, individual-level correlations were good (ρ(2)=0.71, 95% CI 0.71-0.71 for concurrent chemotherapy and ρ(2)=0.61, 0.61-0.61 for modified vs standard radiotherapy) and trial-level correlations very good (R(2)=0.85, 95% CI 0.77-0.92 for concurrent chemotherapy and R(2)=0.95, 0.91-0.98 for modified vs standard radiotherapy). INTERPRETATION: We found a high level of evidence that DFS is a valid surrogate endpoint for overall survival in studies of adjuvant chemotherapy involving patients with non-small-cell lung cancers, and PFS in those of chemotherapy and radiotherapy for patients with locally advanced lung cancers. Extrapolation to targeted agents, however, is not automatically warranted. FUNDING: Programme Hospitalier de Recherche Clinique, Ligue Nationale Contre le Cancer, British Medical Research Council, Sanofi-Aventis.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.013 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".