RESPONSE: Re: Canadian National Breast Screening Study-2: 13-Year Results of a Randomized Trial in Women Aged 50-59 Years
Bibliographic record
Abstract
We thank Professor Narod for his comments. We agree that one possible explanation for our findings (1) is that neither screening with mammography plus physical examination nor physical examination alone was effective. However, breast cancer mortality may have been reduced in both arms of the trial, but that can only be inferred, since it was deemed unethical to have an unscreened control group. The trial design was based on the assumption that mammography screening reduced breast cancer mortality in women aged 50–59 years, and the apparent indicators of mammography effectiveness that Professor Narod lists, cited by others as indicating that breast screening is likely to be effective (2,3), were met in the trial. This suggests an alternative explanation for the coexistence of favorable indicators of mammography effectiveness and the lack of breast cancer mortality reduction compared with physical examination only—namely, that the majority of the small (impalpable) cancers detected by mammography represent pseudo-disease or overdiagnosis. Recently, in the context of lung cancer screening, it has been pointed out that overdiagnosis is being ignored (4). In practice, screening with mammography plus physical examination only slightly reduced the number of large breast cancers (20 mm or more in diameter) compared with physical examination alone during the period of screening in the Canadian National Breast Screening Study-2 [114 and 136, respectively (1)] but not the number of lymph node-positive cancers [92 and 86, respectively (5)]. Thus, the “advanced” breast cancers associated with mortality were not affected by the addition of mammography to physical examination screening, which explains its lack of effect on breast cancer mortality. One reason to doubt overdiagnosis in the context of our trial, is the “catch-up” of invasive breast cancers that occurred in the physical-examination-alone group with continued follow-up. However, after screening ceased in the trial in 1988, most women had access to provincial breast screening programs. It is likely that the majority of our participants, having been sensitized to a possible benefit of breast screening, volunteered for these programs. All of the programs include mammography; thus, the opportunity for overdiagnosis continued for women in both arms of the trial. Therefore, the apparent equivalence of cancers at the end of follow-up does not exclude overdiagnosis as an explanation for many of our findings. Although we agree with Narod that researchers should seek an explanation for our results, we fear that this will not be possible unless they accept the recent proposal that a trial should be initiated comparing screening by use of mammography alone with screening by breast physical examination (6). We urge any researcher interested in such a trial to contact A. B. Miller directly; he will be happy to share a protocol that has been drafted for such a trial. Should such a trial not occur, we can only hope that new biomarkers will eventually result in more effective early detection or that new treatments for breast cancer will be so effective as to make breast screening unnecessary.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.018 | 0.129 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.003 | 0.002 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.030 | 0.019 |
| Insufficient payload (model declined to judge) | 0.031 | 0.012 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".