Film mammography for breast cancer screening in younger women is no longer appropriate because of the demonstrated superiority of digital mammography for this age group
Bibliographic record
Abstract
Arguing against the Proposition is Gary T. Barnes, Ph.D. Dr. Barnes received his B.S. in Physics in 1964 from Case Institute of Technology (now Case-Western Reserve University), Cleveland, Ohio and his Ph.D. in Physics in 1970 from Wayne State University in Detroit. Following completion of his Ph.D., he received postdoctoral training in Medical Physics at the University of Wisconsin, Madison. In 1972 he joined the Department of Radiology at the University of Alabama at Birmingham (UAB) where, from 1976 to 1987, he was Chief of the Department's Physics Section and, from 1987 to 2002, was Director of the Physics and Engineering Division. In 2002 he became Professor Emeritus. He continues to be involved at UAB part time by chairing the Radioisotope and Radiation Safety Committee and teaching radiology residents. He is also involved with X-Ray Imaging Innovations, Inc., a technology development company he founded in 1998. Dr. Barnes has been active on committees of the AAPM, Southeastern Chapter of the AAPM (SEAAPM), ACR, RSNA, and the ABR, and is past President of the AAPM (1988) and SEAAPM (1979). He is a diplomate of the ABR (Radiological Physics), and a Fellow of the AAPM, ACR, and the American Institute of Biomedical Engineers. He was the 2005 recipient of the Coolidge Award. Dr. Barnes’ research interests are in diagnostic x-ray imaging and mammography and include work on scatter control, screen-film systems, digital radiography, and clinical medical physics. He is the author or co-author of 10 patents and scientific papers. When I began working with mammography equipment, it was common for images to be recorded on direct-exposure film. Presumably this was because of the extremely high spatial resolution requirements for detecting breast cancer. Doses typically exceeded , timers were mechanical, proper breast compression was a thing of the future, and automatic exposure control and grids for mammography did not exist. We've come a long way—through xeroradiography and dedicated screen-film systems for mammography—and each technical development has improved the ability to detect breast cancer. Only five years ago, following considerable collaborative development by several academic and industrial groups, digital mammography was introduced commercially. The Digital Mammographic Imaging Screening Trial (DMIST)1 may have been the first time that a new technology for breast imaging was actually put to the acid test to answer the question: “Is the accuracy of digital mammography equal to or higher than that of film?” The conclusion was positive—at least for certain subgroups of women, including those with dense breasts, women under 50, and premenopausal women. Is digital mammography a mature technology? Of course not—there is ample room for further improvement and refinement. There is room for growth of new applications that will have a much greater impact on accuracy than the existence of digital mammography itself-applications like CAD, tomosynthesis, contrast imaging, telemammography, and dual-energy imaging. Each of these applications has the potential to improve the detection, diagnosis, and treatment of breast cancer in different ways. And the use of PACS in conjunction with digital mammography will improve the efficiency of image management. So, until we develop something better, digital mammography should become the norm for breast cancer screening. Women with dense breasts will likely benefit from this technology because of its wide dynamic range, improved DQE, and adjustable brightness and contrast through a user-defined lookup table, as well as through sophisticated image enhancement methods that can be readily applied to digital images. Another thing we learned from DMIST and other studies is that there is enormous inter-reader variability in interpreting mammograms. It is clear that technology is only part of the solution. Just as we have the technology to solve problems like hunger, other factors (like politics and inertia) frequently block our progress. We will gain at least as much in performance of breast imaging by improving and standardizing reader skills as we will in converting from film to digital mammography. What a great idea to do both! The premise for this Point/Counterpoint debate is based on a paper published in the October, 2005 issue of the New England Journal of Medicine.1 A cursory reading of the paper supports the premise. However, the statement (attributed by Mark Twain to Benjamin Disraeli), “There are lies, damn lies and statistics,” is particularly relevant when one reads the paper more carefully. Of interest are the sensitivity and specificity values given in Table 3 of the paper and listed in Table I below. Inescapable in the paper is the authors’ bias for digital mammography. They state: “We found that digital mammography was significantly better than conventional film mammography for detecting breast cancer in young women, premenopausal and perimenopausal women, and women with dense breasts. There was no significant difference in diagnostic accuracy between digital and film mammography in the population as a whole or in other predefined subgroups.” I find it disturbing that the New England Journal of Medicine reviewers and editor allowed the authors to make such blanket and unqualified statements. If there is no difference between digital and film mammography for all women and there is a significant difference in favor of digital mammography for women less than old, one might conclude that film mammography must be better for women greater than old. Although the greater-than--old group was studied, sensitivity and specificity numbers, or for that matter receiver operator characteristic (ROC) curves, were not given. Even though sensitivity and specificity values are presented, the majority of the authors’ conclusions are based on comparing areas under ROC curves. Comparing ROC curve areas gives excessive weighting to regions that have high false positive rates that are not relevant for screening. For example, the authors conclude that digital is significantly better than film mammography for dense breasts based on ROC area comparisons. This conclusion is not supported by the sensitivity and specificity values presented in the paper. Another concern is that the paper compares the performance of digital mammography units with existing, and in many instances old and poor performing, film mammography units. The digital units were all new and tweaked to perform at the highest level. The authors looked for differences in the units of the four digital-mammography manufacturers that were included in the study (none were seen), but lumped all film mammography units together. As a reader of accreditation phantom images for the American College of Radiology (ACR) Mammography Accreditation Program (MAP) I, and other reviewers have observed that there are very noticeable differences in image quality of sites that have been accredited by the ACR. There are some mammography units with better scatter control that consistently yield better accreditation phantom images than those obtained with other units. The quality of screen-film processing impacts accreditation phantom and clinical images. The results of the study may well have been different if only state-of-the-art film mammography units with good scatter control and good screen-film processing had been compared with digital units. In summary, the premise that film mammography is inferior to digital mammography and should not be used in examining young women is built on sand, not stone. The authors of the paper on which the premise is based are biased. They have chosen the statistics that support their conclusions, and they do not adequately discuss other statistics that are less supportive of their bias. It is not clear that the claimed improved diagnostic performance with digital would be realized if a site has a film mammography unit with good scatter control and good screen-film processing. I am disappointed that Dr. Barnes has used one of the classic logical fallacies, the “ad hominem” argument2—i.e., when you aren't able to attack a position held by a person, attack the person. He suggests that the DMIST researchers were biased. The DMIST results were published under peer review in a highly reputable journal. Is he also suggesting that the reviewers were biased? What is the evidence? In fact, a respected, independent statistical group blinded the researchers to the data until completion of the analysis. Until the investigators saw these results, they expressed concerns that a difference between the modalities would not be detectible because of the large reader variability. While there was hope that digital mammography might prove to be advantageous, there was a healthy level of skepticism as well. Furthermore, while the initial policy was to attempt to match doses between the two modalities, the investigators allowed the doses from digital to drop during the trial based on evidence from physics testing. If anything, this would work against an attempt to make digital mammography look better than film. Dr. Barnes also suggests that the digital systems operated in peak form while the film systems were suboptimal. Again he provides no evidence for this claim. In fact, study sites were selected for their commitment to high quality and all imaging systems were MQSA compliant. The performance of both digital and film mammography was carefully monitored. In the case of sites using photostimulable phosphor systems, film and digital mammography was performed on the same units. Did the local medical physicists and technologists at all 33 sites work to sabotage film in comparison with digital? Film mammography is a mature technique, while the steep learning curve for both manufacturers and radiologists in understanding how best to use the new technology probably was a factor that could have undermined the success of digital mammography in DMIST. Probably of greatest concern, however, is Dr. Barnes’ inference that because there was no statistically significant benefit of digital mammography across the entire study population, while there was proven benefit for younger women and those with dense breasts, implies that film must have performed better in older women or those with fatty breasts. This is simply not the case. If one thinks about the nature of a statistical test, it involves comparing the difference between two values (in this case, areas under an ROC curve) to the statistical uncertainty in that difference. The analogy in imaging theory is Rose's criterion for object detectability of a signal difference in the presence of noise.3 In DMIST, it was found that the difference was significant for younger women and those with denser breasts. While the difference was also positive (i.e., consistent with digital mammography being superior) in the overall population, the noise in this difference caused the results to be insignificant. There was no magic and no deception. Furthermore, the investigators did not “cherry pick” the data to obtain a result that was significant. The analyses were preplanned before the study was started. Limitations on article length precluded complete publication of all analyses, but a full subset analysis will be published shortly. Finally, Dr. Barnes complains that the areas under the ROC curves were computed using the entire curve (i.e., using the standard approach). He suggests that this approach weighs the results inappropriately towards regions of the curve representing low specificity. Although it is conceivable that some other method might be more relevant, I'm not aware of a better method that has been validated. I am encouraged that the ROC curves show clear separation in favor of digital even at low false positive fractions typical of the use of mammography in screening, and that the curves show no indication of crossing at any specificity level. Digital mammography was developed primarily to address limitations of film for imaging women with dense breasts, frequently younger women. In DMIST, of the 165 cancers in women with dense breasts, 24% were found only on digital compared with 11.5% found only on film.1 We should be delighted that mammography can now be more helpful for those women. I agree with my esteemed colleague that mammography has improved markedly in the past . I also agree that digital mammography is not a mature technology. I disagree that digital should be the norm for screening. As noted in my Opening Statement, Reference 1 is biased. It is not clear that digital is superior to state-of-the-art film mammography. Furthermore, digital is more expensive. The cost of a digital unit and review workstation is four or five times that of a conventional film unit and processor. The greater patient throughput of digital units is compromised by downtime at the sites, at least those I support. Service contracts are expensive and more than offset the cost of service and of film and chemistry consumed in film mammography. One inefficiency of digital is that it takes longer for radiologists to read an exam compared with film. We have a shortage of radiologists reading mammography, and digital aggravates the problem. For these reasons I suspect that a cost-benefit analysis (based on good data and science) would show that state-of-the-art film mammography is superior to digital mammography. A travesty is that Centers for Medicare & Medicaid Services (CMS) reimburses more for digital than for film mammography. Reimbursing more for the same study and diagnostic performance just because the equipment costs more is an absurdity and is caused by lobbying by manufacturers. Digital mammography is a factor contributing to the spiraling costs of health care with no benefit. It is disturbing how poor mammography is, either digital or film, at detecting cancer. A sensitivity of is not good. We can do better. Based on our experience with CT and the ability of 3D x-ray techniques to remove superimposed overlying and underlying structures and improve lesion conspicuity, it is my opinion that 3D x-ray tomosynthesis or cone beam CT techniques are the future of breast imaging. These will be available in . In addition to better sensitivity, it is also likely that 3D x-ray imaging will have better CAD performance (computers as well as humans do better with simpler images). In conclusion, it is not prudent to invest in a digital mammography system today. If one does, within a few years it will be necessary to spend a comparable sum once again to purchase a 3D breast x-ray unit. Be wary of salesmen who promise that a digital unit can be upgraded without documenting the cost. Two-dimensional projection mammography, whether digital or film, is nearing the end of its tenure. Wait a couple of years and buy a 3D x-ray unit. It will be a better investment and worth the wait.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".