Correlation between phase 2 clinical trial design and subsequent phase 3 outcome.
Bibliographic record
Abstract
2549 Background: Randomized phase 2 (RP2) trials have been postulated to better predict phase 3 (P3) outcome than single arm trials. The aim of this study was to determine which characteristics of P2 trials that led to the conduct of P3 trials are associated with a positive P3 outcome. Methods: Randomized P3 trials testing systemic therapy in patients with malignancy, published in English from 2007 to 2012, were identified through a Medline search. Those of superiority design which cited P2 trials contributing to development of the experimental arm were included. Detailed information regarding design and outcome of P2 and P3 trials was extracted in duplicate. Discrepancies were adjudicated by a third investigator. Positivity was defined as statistical difference between arms for primary endpoint in P3. P2 trials reporting that a pre-stated numeric or statistical target had been met were deemed positive. Statistical analysis was performed using the Generalized Estimating Equation model correlating P2 features with P3 outcome, accounting for any P3 duplication. Results: Of 189 eligible P3 trials 19% were in hematologic malignancies and 81% in solid tumors. The primary outcome was positive in 79 (42%). These were supported by 336 P2 trials (range 1 - 9 per P3 trial); 66 of which were RP2s (29 selection, 21 control, 3 discontinuation, 7 phase 2/3 and a combination of these in 6). RP2 trials were not associated with P3 positivity (see Table, p = 0.44). Positive P2 outcome correlated with positive P3 outcome (p = 0.031). P2 trial features found not to be predictive of P3 outcome included primary endpoint, sponsorship, sample size, similarity in patient population and therapy. Conclusions: We did not find RP2 more predictive than single arm trials. When designing P2 trials consideration should be given to the increased resources required to conduct RP2 studies and anticipated gain for a given disease. 336 phase 2 trials 66 randomized phase 2 270 single-arm phase 2 35 positive 31 not positive* 102 positive 168 not positive* P3 positive P3 negative P3 positive P3 negative P3 positive P3 negative P3 positive P3 negative 14 (40%) 21 (60%) 8 (26%) 23 (74%) 50 (49%) 52 (51%) 61 (36%) 107 (64%) * Numeric or statistical target not stated or not met.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.126 | 0.666 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.005 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".