Magnitude of the Benefit of Progression-Free Survival as a Potential Surrogate Marker in Phase 3 Trials Assessing Targeted Agents in Molecularly Selected Patients with Advanced Non-Small Cell Lung Cancer: Systematic Review
Bibliographic record
Abstract
BACKGROUND: In evaluation of the clinical benefit of a new targeted agent in a phase 3 trial enrolling molecularly selected patients with advanced non-small cell lung cancer (NSCLC), overall survival (OS) as an endpoint seems to be of limited use because of a high level of treatment crossover for ethical reasons. A more efficient and useful indicator for assessing efficacy is needed. METHODS AND FINDINGS: We identified 18 phase 3 trials in the literature investigating EGFR-tyrosine kinase inhibitor (TKIs) or ALK-TKIs, now approved for use to treat NSCLC, compared with standard cytotoxic chemotherapy (eight trials were performed in molecularly selected patients and ten using an "all-comer" design). Receiver operating characteristic analysis was used to identify the best threshold by which to divide the groups. Although trials enrolling molecularly selected patients and all-comer trials had similar OS-hazard ratios (OS-HRs) (0.99 vs. 1.04), the former exhibited greater progression-free survival-hazard ratios (PFS-HR) (mean, 0.40 vs. 1.01; P<0.01). A PFS-HR of 0.60 successfully distinguished between the two types of trials (sensitivity 100%, specificity 100%). The odds ratio for overall response was higher in trials with molecularly selected patients than in all-comer trials (mean: 6.10 vs. 1.64; P<0.01). An odds ratio of 3.40 for response afforded a sensitivity of 88% and a specificity of 90%. CONCLUSION: The notably enhanced PFS benefit was quite specific to trials with molecularly selected patients. A PFS-HR cutoff of ∼0.6 may help detect clinical benefit of molecular targeted agents in which OS is of limited use, although desired threshold might differ in an individual trial.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.034 | 0.123 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.012 | 0.012 |
| Bibliometrics | 0.006 | 0.007 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".