Long-Term Outcomes and Cost-Effectiveness of Breast Cancer Screening With Digital Breast Tomosynthesis in the United States
Bibliographic record
Abstract
BACKGROUND: Digital breast tomosynthesis (DBT) is increasingly being used for routine breast cancer screening. We projected the long-term impact and cost-effectiveness of DBT compared to conventional digital mammography (DM) for breast cancer screening in the United States. METHODS: Three Cancer Intervention and Surveillance Modeling Network breast cancer models simulated US women ages 40 years and older undergoing breast cancer screening with either DBT or DM starting in 2011 and continuing for the lifetime of the cohort. Screening performance estimates were based on observational data; in an alternative scenario, we assumed 4% higher sensitivity for DBT. Analyses used federal payer perspective; costs and utilities were discounted at 3% annually. Outcomes included breast cancer deaths, quality-adjusted life-years (QALYs), false-positive examinations, costs, and incremental cost-effectiveness ratios (ICERs). RESULTS: Compared to DM, DBT screening resulted in a slight reduction in breast cancer deaths (range across models 0-0.21 per 1000 women), small increase in QALYs (1.97-3.27 per 1000 women), and a 24-28% reduction in false-positive exams (237-268 per 1000 women) relative to DM. ICERs ranged from $195 026 to $270 135 per QALY for DBT relative to DM. When assuming 4% higher DBT sensitivity, ICERs decreased to $130 533-$156 624 per QALY. ICERs were sensitive to DBT costs, decreasing to $78 731 to $168 883 and $52 918 to $118 048 when the additional cost of DBT was reduced to $36 and $26 (from baseline of $56), respectively. CONCLUSION: DBT reduces false-positive exams while achieving similar or slightly improved health benefits. At current reimbursement rates, the additional costs of DBT screening are likely high relative to the benefits gained; however, DBT could be cost-effective at lower screening costs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".