Use and Costs of Disease Monitoring in Women With Metastatic Breast Cancer
Bibliographic record
Abstract
PURPOSE: The optimal frequency of monitoring patients with metastatic breast cancer (MBC) is unknown; however, data suggest that intensive monitoring does not improve outcomes. We performed a population-based analysis to evaluate patterns and predictors of extreme use of disease-monitoring tests (serum tumor markers [STMs] and radiographic imaging) among women with MBC. METHODS: The SEER-Medicare database was used to identify women with MBC diagnosed from 2002 to 2011 who underwent disease monitoring. Billing dates of STMs (carcinoembryonic antigen and/or cancer antigen 15-3/cancer antigen 27.29) and imaging tests (computed tomography and/or positron emission tomography) were recorded; if more than one STM or imaging test were completed on the same day, they were counted once. We defined extreme use as > 12 STM and/or more than four radiographic imaging tests in a 12-month period. Multivariable analysis was used to identify factors associated with extreme use. In extreme users, total health care costs and end-of-life health care utilization were compared with the rest of the study population. RESULTS: We identified 2,460 eligible patients. Of these, 924 (37.6%) were extreme users of disease-monitoring tests. Factors significantly associated with extreme use were hormone receptor-negative MBC (odds ratio [OR], 1.63; 95% CI, 1.27 to 2.08), history of a positron emission tomography scan (OR, 2.92; 95% CI, 2.40 to 3.55), and more frequent oncology office visits (OR, 3.14; 95% CI, 2.49 to 3.96). Medical costs per year were 59.2% higher in extreme users. Extreme users were more likely to use emergency department and hospice services at the end of life. CONCLUSION: Despite an unknown clinical benefit, approximately one third of elderly women with MBC were extreme users of disease-monitoring tests. Higher use of disease-monitoring tests was associated with higher total health care costs. Efforts to understand the optimal frequency of monitoring are needed to inform clinical practice.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".