Trends in the design and interpretation of metastatic colorectal cancer phase III clinical trials.
Bibliographic record
Abstract
692 Background: Increasing use of subsequent lines of therapy and crossover in phase III randomized clinical trials (P3 RCTs) has shifted how we perceive the effectiveness of treatments for metastatic colorectal cancer (mCRC). This study aims to characterize the evolution of P3 RCTs in mCRC with respect to clinical trial design and result interpretation. Methods: Abstracts of P3 RCTs of systemic therapy for mCRC conducted between 1980 and 2014 were identified by searching PubMed, Medline, and ASCO abstracts. Data regarding trial design, agent(s) investigated, primary endpoint, secondary endpoint(s), primary endpoint significance and interpretation of the study results (conclusions) were extracted. Results: A total of 422 trials were identified by the search strategy, and 132 eligible trials were included. Over time the sample size of P3 RCTs in mCRC has been increasing and there has been a steady increase in trials studying targeted therapy (see table below for detailed results by decade). A trend towards a smaller percentage of P3 RCTs sponsored by co-operative groups has been observed in recent decades. The most common primary endpoint was overall survival (OS) which was used in 35% of the trials. A decreasing trend in the use of OS was observed since the 1990s. Other common primary endpoints include: progression-free survival (PFS) in 28% and response rate (RR) in 20% of the P3 RCTs. The primary endpoint was met in 45% of the trials. There was discordance between the primary endpoint significance and the authors’ conclusions in 14% of the trials. Conclusions: The design and interpretation of P3 RCTs for mCRC has changed over time from 1980 to present. The use of OS as the primary endpoint is decreasing, while the use of PFS is increasing. [Table: see text]
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.770 | 0.875 |
| Meta-epidemiology (narrow) | 0.002 | 0.003 |
| Meta-epidemiology (broad) | 0.006 | 0.007 |
| Bibliometrics | 0.018 | 0.025 |
| Science and technology studies | 0.001 | 0.006 |
| Scholarly communication | 0.017 | 0.012 |
| Open science | 0.005 | 0.006 |
| Research integrity | 0.006 | 0.008 |
| Insufficient payload (model declined to judge) | 0.004 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".