RE: Advanced Breast Cancer Definitions by Staging System Examined in the Breast Cancer Surveillance Consortium
Bibliographic record
Abstract
As investigators for ECOG-ACRIN’s Tomosynthesis Mammographic Imaging Screening Trial (TMIST) trial, we are writing to draw attention to conceptual issues in the outcome definitions and study population in Kerlikowske et al. (1), which limit inferences with respect to the TMIST trial. Kerlikowske et al. (1) converted the TMIST primary outcome definition into a staging system for breast cancer and compared it with other staging systems in association with 5-year breast cancer mortality. However, the primary outcome of TMIST is not a cancer staging system but simply a binary classification of cancers as “advanced” or not. TMIST’s endpoint of advanced cancers was defined to identify cancers that generally require chemotherapy, because although chemotherapy prevents many cancer-related deaths, it is also associated with clinically significant morbidity. Reducing chemotherapy-related morbidity is a valuable goal of breast cancer screening. The authors constructed an ordinal categorical response using elements of the TMIST binary endpoint and performed Receiver Operating Characteristic (ROC) analysis on this ordinal categorical response (1). Although ROC analysis cannot be performed as a binary outcome, the relevance of the ordinal comparison for TMIST is not clear. A more relevant comparison would be conducted with the binary assessment that would result from using an American Joint Committee on Cancer stage as threshold for advanced cancer. For example, if stage IIA or IIB is used as the threshold, as was done by Kerlikowske et al. (1), one can estimate measures of performance that are appropriate for binary tests. The relevant measures for predicting cancer death in 5 years, given at the bottom of Table 2 in the JNCI article for the American Joint Committee on Cancer staging systems and at the bottom of Table 3 for the TMIST definition (1), are combined in Table 1 here. Measures for predicting cancer death in 5 years by AJCC staging systems and by TMIST definitiona AJCC = American Joint Committee on Cancer; AJCC Anat = American Joint Committee on Cancer Anatomic stage; AJCC Progn = American Joint Committee on Cancer Prognostic Pathologic stage; TMIST = Tomosynthesis Mammographic Imaging Screening Trial. Measures for predicting cancer death in 5 years by AJCC staging systems and by TMIST definitiona AJCC = American Joint Committee on Cancer; AJCC Anat = American Joint Committee on Cancer Anatomic stage; AJCC Progn = American Joint Committee on Cancer Prognostic Pathologic stage; TMIST = Tomosynthesis Mammographic Imaging Screening Trial. Another important difference in the outcomes relates to follow-up. The article considers 5-year risk of death, which overrepresents deaths from Estrogen Receptor (ER)-negative cancer and neglects longer term risk of ER+ deaths. The majority of screen-detected breast cancers are ER+, and it is important to address mortality from these cancers. The 2-county trial in Sweden showed that more than 15 years of follow-up was needed to demonstrate the full mortality reduction of breast cancer screening and showed that even at 10 years, fewer than one-half of the averted deaths had been observed (2-4). Finally, Kerlikowske et al. (1) report a large (approximately 60%) proportion of advanced cancer in the Breast Cancer Surveillance Consortium (BCSC) population (Table 3), underscoring that the study population was probably not a pure screening population and likely includes symptomatic women, as commonly seen in practice-based (nontrial) data (5). These important conceptual differences limit the implications of Kerlikowske et al. (1) for TMIST. TMIST is conducted by the ECOG-ACRIN Cancer Research Group (Peter J. O’Dwyer, MD, and Mitchell D. Schnall, MD, PhD, Group Co-Chairs) and supported by the National Cancer Institute of the National Institutes of Health (NIH) (award number: UG1CA189828). Role of the funder: The funder had no role in the writing of the correspondence or decision to submit it for publication. Disclosures: The authors all receive funding from ECOG-ACRIN for their work, but have no other disclosures. Author contributions: Conceptualization: EDP, CG, JS, MAT, MY, MDS. Data Curation: CG. Formal Analysis: CG, MAT, MY. Funding Acquisition: EDP, CG, EC, MDS. Investigation: MAT, MY. Methodology: EDP, CG, MAT, MY, EC, MDS. Project Administration: EDP, CG, MAT, MY, EC, MDS. Resources: EDP, CG, MAT, MY, EC. Software: CG, MY. Supervision: EDP, CG, MAT, MY, EC. Validation: CG. Visualization: CG. Writing, original draft: EDP, CG. Writing, review and edit: EDP, CG, JS, MAT, MY, EC, MDS. Disclaimer: The content is solely the opinion of the authors and does not necessarily represent the views of the NIH. The data underlying this correspondence are available in the correspondence itself.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.031 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.004 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.026 | 0.019 |
| Insufficient payload (model declined to judge) | 0.017 | 0.018 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".