Impact of the 2010 Consensus Recommendations of the Clinical Trial Design Task Force of the NCI Investigational Drug Steering Committee
Bibliographic record
Abstract
Oncology phase III trials have a high failure rate, leading to high development costs. The Clinical Trials Design Task Force of the Investigational Drug Steering Committee of the NCI Cancer Therapy and Evaluation Program developed Recommendations regarding the design of phase II trials. We report here on the results of a Concordance Group review charged with documenting whether concordance rates improved after the publication of the Recommendations. One hundred and fifty-five trials were reviewed. Letter of Intents (LOI) from the post-Recommendation period were more likely to be randomized (44% vs. 34%) and biomarker selected (19% vs. 10%). Single-arm studies using time-to-event endpoints (benchmarked against historical data) were similar, as was the type of tumor. There was a significant improvement in the rate of concordance, with 74% of LOIs scored as concordant compared with 58% before the Recommendations (P = 0.042). This included a marked decrease in the use of single-arm designs to evaluate the activity of drug combinations (19% vs. 5%, P = 0.009). There were areas for which clarification was warranted, including the need for protocols to include further development plans, the use of realistic benchmarks, the careful evaluation of historical controls, and the use of a standard treatment option as a control. Ongoing critical evaluation of current trial design methodology and the development of new Guidelines when appropriate will continue to improve drug development ensuring that safe and effective cancer therapeutics are made available to our patients as quickly and efficiently as possible.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.068 | 0.521 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.006 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".