The Ethics of Early Stopping Rules
Bibliographic record
Abstract
2 and as chair of the Data Safety Monitoring Committee for one of the trials (the Intergroup MA17 trial of letrozole) he criticizes, we are concerned that this article has the potential to impair significantly the conduct of future trials in breast cancer adjuvant endocrine therapy and consequently be harmful to patients.Our major issues with the article are: (1) Dr Cannistra repeats some of the misunderstandings about MA17 that we have previously addressed.3 Most important in this context is the implication that disease-free survival (DFS) was not the sole primary end point around which the study was designed and that the difference in DFS targeted was at a specific point in time, rather than, as was actually the case, a hazard ratio measured over time.(2) A one-size-fits-all approach to choosing end points in cancer clinical trials is wrong.The proper choice of end point is context (disease, stage, and therapy) -specific.DFS is an appropriate outcome in trials of endocrine therapy in early breast cancer.(3) The title and tone of the article implies that stopping MA17 "early" meant the goals of the trial had not been achieved because of the intervention of a Data Safety Monitoring Board (DSMB).On the contrary, the study question was answered sooner than we expected because the treatment effect was larger than anticipated.(4) The value of Dr Cannistra's comments and recommendations regarding DSMB's are undermined by an apparent lack of familiarity with how these bodies function under the auspices of a National Cancer Institute cooperative group program.Because of their importance, we will expand briefly on points 2 and 4 in the following paragraphs: 2. Disease-free survival is commonly chosen as the end point of US Food and Drug Administration-and NCIapproved trials of adjuvant endocrine therapy in breast cancer.We feel strongly that this is appropriate, not only because the accumulated experience to date suggests that DFS is a surrogate for overall survival in this setting, 4 but also because preventing breast cancer recurrence has value in itself.Women who suffer an in-breast recurrence often require the mastectomy their initial management was intended to avoid; women who develop new breast cancer repeat the trauma of their initial diagnosis and treatment; and women with distant recurrence inevitably die of their disease.Even in the absence of data that formally quantify
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.180 | 0.432 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.003 | 0.015 |
| Scholarly communication | 0.008 | 0.005 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.035 | 0.054 |
| Insufficient payload (model declined to judge) | 0.003 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".