Bibliographic record
Abstract
Figure: Harmon J. EyreFigure: Robert A. SmithLate in October 2001, The Lancet published a research letter by Ole Olsen and Peter Gøtzsche,1 two researchers from the Nordic Cochrane Centre, summarizing the results of their analysis and critique of the seven randomized trials of breast cancer screening. The authors concluded that there was no reliable evidence that screening for breast cancer reduces mortality and therefore was unjustified, a conclusion they had reached in an earlier analysis, also published in Lancet.2 Furthermore, they concluded that screening leads to more aggressive treatment, and thus results in greater harms than not screening at all. These conclusions are startling, since they are at variance with the most fundamental understanding of breast cancer as a progressive disease, as well as decades of supporting scientific evidence from the individual trials, meta-analyses, observational studies, and confirmatory, independent expert reviews conducted by medical and scientific groups in North American and Europe. The Lancet publication was followed by considerable media attention, with some quoted experts giving credence to the claims of Olsen and Gøtzsche, while others dismissed the report as deeply flawed and without value. While leading organizations came together to reassure women and their doctors that the conclusions of the Cochrane Report were not credible, the National Cancer Institute's PDQ group announced that they concurred with the conclusions of Olsen and Gøtzsche and would revise the current statement on mammography screening on the NCI Web site to indicate the lack of confirmatory scientific evidence. This statement was followed by a press release from the NCI stating that the PDQ group was advisory to the NCI, but did not set policy, and that the NCI would stand by its guideline that women in their 40s should begin regular screening for breast cancer with mammography every one to two years. The presence of conflicting opinions from scientists and institutions inevitably leads to the question, “Could we have been that wrong about the beneficial effects of mammography?” The answer is No. It is important at the outset to know the methodological underpinnings of Olsen and Gøtzsche's conclusions. The review is not an analysis of new data, but rather a re-analysis of data that have been carefully scrutinized repeatedly by independent expert groups representing many professional organizations and government agencies in the United States and Europe. Consistently and overwhelmingly, independent analysis and re-analysis of individual studies, as well as expert reviews of the totality of the data, have led to the conclusion that mammography is beneficial. It is beneficial because it advances the time of diagnosis, and the curative potential of a small breast tumor is greater than that of a large breast tumor. However, Olsen and Gøtzsche assert that none of the randomized trials of breast cancer screening were of high quality, and only two of seven were of medium quality (the Malmö trial and the Canadian trial). The most common factor for their conclusion that the remaining five trials (Two-county, Stockholm, Göteborg, Edinburgh, and Health Insurance Plan of New York) were of poor or flawed quality was the randomization process, the manner in which end-result committees determined cause of death, and the lack of difference in all-cause mortality between experimental and control groups. These judgments about study quality are at the core of their methodology, and have been vigorously contested by screening experts. Very large trials invariably require logistic compromises not required of smaller, hospital-based therapy trials. Complex statistical routines are employed to insure data are adjusted for randomization and sampling strategies, as well as imbalances between groups that are observed during the course of the study. Olsen and Gøtzsche have tended to highlight as flaws study characteristics that have been shown to be inconsequential and have no effect on end results. Of the remaining two trials judged to be of acceptable quality, the Malmö trial end results puzzlingly are from a 1987 report showing no benefit, rather than the more recent report showing a statistically significant 19% reduction in breast cancer deaths. The Canadian trial also showed no benefit from screening, although it is ironic that it should be held up as an example of higher quality. Numerous publications have identified significant problems with mammographic image quality and the randomization process, factors that likely account for the excess rate of breast cancer deaths in the experimental group at the conclusion of the study. At best, we can judge Olsen and Gøtzsche's assessment of study quality as inconsistent. However, when you eliminate studies from your meta-analysis that show a benefit from screening, and leave only studies that showed no benefit, it is clear that you will ultimately observe no benefit. Thus, their conclusion is not surprising. Olsen and Gøtzsche also argue that breast cancer mortality is an inappropriate endpoint, since cause of death committees invariably tend to attribute deaths in the control group to breast cancer, and deaths in the experimental group to some cause other than breast cancer. Rather, they contend, either all-cause-mortality or all-cause-cancer-mortality is a more appropriate endpoint for analysis. Setting aside the fact that errors in attributing cause of death would have to be so significantly systematic to overcome the underestimate of benefit in trial (due to non-compliance with randomization assignment), this argument also is unjustified on numerous other methodological points. First, an intervention focused on reducing premature mortality from breast cancer cannot be expected to contribute to mortality reductions from heart disease, diabetes, colorectal cancer, or trauma. Second, if breast cancer accounts for three to four percent of all deaths in women, a 30–60% reduction in breast cancer mortality would be observed as a 1–2% reduction in all-cause mortality, virtually impossible to observe in the average-sized trial. Finally, the authors assert that screening leads to more aggressive treatment, a bewildering conclusion since rates of breast-conserving therapy have increased as a consequence of the opportunity to treat small tumors less aggressively. This conclusion is based on other meta-analyses of early trial data showing higher rates of mastectomy, and radiation therapy studies showing an excess of cardiac deaths in women receiving radiation to the axilla and tumor bed. More specifically, the findings are based on old data, representing clinical practice 25 to 30 years ago, including archaic approaches to radiation therapy prior to the appreciation and opportunity to avoid the heart with better field design. The research letter was accompanied by a commentary from Lancet Editor Richard Horton, MD,3 in which he agreed with the conclusions and commended the authors for their systematic review, while also assailing the editors of the Cochrane Breast Cancer Group for disagreeing with the authors' methods and conclusions. The Cochrane editors had insisted that a more balanced statement about screening be added to the final report, and that publication of the conclusion that screening leads to more aggressive treatment be postponed pending further deliberation. Despite the weight of the evidence supporting the benefits of early detection, how is it that some in the scientific community and media are so quick to embrace the conclusions of such a flawed report? Setting aside the sensational elements of this story, taking sides in this debate can be influenced by fundamental opinions about the value of screening per se, and discontent with what is regarded as excessive financial and human costs associated with screening. Others have expressed a fundamental dissatisfaction with early detection as a final or interim disease-control strategy, and regard any resources devoted to screening as both an expression of acquiescence to an incomplete solution, as well as resources diverted from an investment into prevention research. While experts and advocates can legitimately hold different opinions about the value of screening and cost effectiveness, it also is natural that they would seek to either reinforce those views, or allow them to be secondary to so-called evidence alleging that there is no supporting scientific evidence for screening policy in the first place. As it stands, not only is there little constructive criticism in the current debate, since the claims of Olsen and Gøtzsche are not sustainable by the evidence, but greater harms may result from women and physicians having their confidence in mammography shaken, even if only momentarily. Breast cancer screening with mammography has been evaluated with millions of person-years of experimental evidence which has shown a substantial and significant reduction in breast cancer mortality associated with an invitation to screening. More recently, analysis of cumulative breast cancer mortality in two Swedish counties over a 29-year period showed that the policy of offering breast cancer screening to the population reduced overall breast cancer mortality by 50 percent. Among women actually attending screening, the risk of dying from breast cancer was reduced by 63 percent compared with women who did not get screened.4 In the US, breast cancer mortality rates began to decline for the first time in 1989, and declined by about two percent per year until 1995, and have declined by 3.4 percent since 1995–1998. The declines in mortality are consistent with the trend towards diagnosis of breast cancer at more favorable stages, and a decline in the incidence rate of advanced disease. Improvements in awareness, early detection, and therapy are resulting in a decline in mortality from this disease. Health care professionals can be confident in the prevailing recommendations supporting routine breast cancer screening that have been issued by leading professional societies and national governments. These recommendations derive from numerous periodic independent reviews conducted by expert groups, each of which has consistently reaffirmed the fundamental conclusion that early breast cancer detection and treatment results in decreased breast cancer mortality. Clinicians also should recognize that their patients look to them as the ultimate authority for guidance and direction, and that they must take the opportunity to assure their patients that the scientific evidence supporting the benefit of early breast cancer detection is sound and in no way challenged by this recent report.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".