Meta-Analysis of Phase II Cooperative Group Trials in Metastatic Stage IV Melanoma to Determine Progression-Free and Overall Survival Benchmarks for Future Phase II Trials
Bibliographic record
Abstract
PURPOSE: Objective tumor response rates observed in phase II trials for metastatic melanoma have historically not provided a reliable indicator of meaningful survival benefits. To facilitate using overall survival (OS) or progression-free survival (PFS) as an endpoint for future phase II trials, we evaluated historical data from cooperative group phase II trials to attempt to develop benchmarks for OS and PFS as reference points for future phase II trials. PATIENTS AND METHODS: Individual-level and trial-level data were obtained for patients enrolled onto 42 phase II trials (70 trial arms) that completed accrual in the years 1975 through 2005 and conducted by Southwest Oncology Group, Eastern Cooperative Oncology Group, Cancer and Leukemia Group B, North Central Cancer Treatment Group, and the Clinical Trials Group of the National Cancer Institute of Canada. Univariate and multivariate analyses were performed to identify prognostic variables, and between-trial(-arm) variability in 1-year OS rates and 6-month PFS rates were examined. RESULTS: Statistically significant individual-level and trial-level prognostic factors found in a multivariate survival analysis for OS were performance status, presence of visceral disease, sex, and whether the trial excluded patients with brain metastases. Performance status, sex, and age were statistically significant prognostic factors for PFS. Controlling for these prognostic variables essentially eliminated between-trial variability in 1-year OS rates but not in 6-month PFS rates. CONCLUSION: Benchmarks are provided for 1-year OS or OS curves that make use of the distribution of prognostic factors of the patients in the phase II trial. A similar benchmark for 6-month PFS is provided, but its use is more problematic because of residual between-trial variation in this endpoint.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.155 | 0.195 |
| Meta-epidemiology (narrow) | 0.004 | 0.001 |
| Meta-epidemiology (broad) | 0.018 | 0.039 |
| Bibliometrics | 0.007 | 0.006 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.005 | 0.003 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".