Updated Standardized Definitions for Efficacy End Points (STEEP) in Adjuvant Breast Cancer Clinical Trials: STEEP Version 2.0
Bibliographic record
Abstract
PURPOSE: The Standardized Definitions for Efficacy End Points (STEEP) criteria, established in 2007, provide standardized definitions of adjuvant breast cancer clinical trial end points. Given the evolution of breast cancer clinical trials and improvements in outcomes, a panel of experts reviewed the STEEP criteria to determine whether modifications are needed. METHODS: We conducted systematic searches of ClinicalTrials.gov for adjuvant systemic and local-regional therapy trials for breast cancer to investigate if the primary end points reported met STEEP criteria. On the basis of common STEEP deviations, we performed a series of simulations to evaluate the effect of excluding non-breast cancer deaths and new nonbreast primary cancers from the invasive disease-free survival end point. RESULTS: Among 11 phase III breast cancer trials with primary efficacy end points, three had primary end points that followed STEEP criteria, four used STEEP definitions but not the corresponding end point names, and four used end points that were not included in the original STEEP manuscript. Simulation modeling demonstrated that inclusion of second nonbreast primary cancer can increase the probability of incorrect inferences, can decrease power to detect clinically relevant efficacy effects, and may mask differences in recurrence rates, especially when recurrence rates are low. CONCLUSION: We recommend an additional end point, invasive breast cancer-free survival, which includes all invasive disease-free survival events except second nonbreast primary cancers. This end point should be considered for trials in which the toxicities of agents are well-known and where the risk of second primary cancer is small. Additionally, we provide end point recommendations for local therapy trials, low-risk populations, noninferiority trials, and trials incorporating patient-reported outcomes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.292 | 0.505 |
| Meta-epidemiology (narrow) | 0.004 | 0.003 |
| Meta-epidemiology (broad) | 0.007 | 0.017 |
| Bibliometrics | 0.013 | 0.014 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.010 | 0.006 |
| Open science | 0.006 | 0.009 |
| Research integrity | 0.005 | 0.013 |
| Insufficient payload (model declined to judge) | 0.011 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".