Conclusions for mammography screening after 25-year follow-up of the Canadian National Breast Cancer Screening Study (CNBSS)
Bibliographic record
Abstract
UNLABELLED: Twenty-five-year follow-up data of the Canadian National Breast Cancer Screening Study (CNBSS) indicated no mortality reduction. What conclusions should be drawn? After conducting a systematic literature search and narrative analysis, we wish to recapitulate important details of this study, which may have been neglected: Sixty-eight percent of all included cancers were palpable, a situation that does not allow testing the value of early detection. Randomisation was performed at the sites after palpation, while blinding was not guaranteed. In the first round, this "randomisation" assigned 19/24 late stage cancers to the mammography group and only five to the control group, supporting the suspicion of severe errors in the randomisation process. The responsible physicist rated mammography quality as "far below state of the art of that time". Radiological advisors resigned during the study due to unacceptable image quality, training, and medical quality assurance. Each described problem may strongly influence the results between study and control groups. Twenty-five years of follow-up cannot heal these fundamental problems. This study is inappropriate for evidence-based conclusions. The technology and quality assurance of the diagnostic chain is shown to be contrary to today's screening programmes, and the results of the CNBSS are not applicable to them. KEY POINTS: • The evidence base of the Canadian study (CNBSS) has to be questioned.• Severe flaws in the randomization process and test methods occurred. • Problems were criticized during and after conclusion of the trial by experts.• The results are not applicable to quality-assured screening programs. • The evidence base of this study must be re-analyzed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".