Predicting Expenditures for Persons With Chronic Conditions
Bibliographic record
Abstract
Research on the distribution of health care expenditures among the U.S. population has shown that a small proportion of individuals, many of whom have chronic health conditions, accounts for a disproportionate share of total expenses. Previous work by Berk and Monheit showed that in both 1987 and 1996 the top 1% of individuals in the expenditure distribution accounted for more than one-quarter of all expenses, and the top 5% accounted for more than half. They found this distribution had been constant for more than 20 years, despite significant changes in the U.S. health care system during that period. Previous research has also shown that more than three-quarters of all medical expenditures are associated with the care of persons with chronic conditions. Recent efforts to contain medical care costs in the U.S. have focused on what might be done to obviate some of this disproportionate spending, either through more effective use of preventive care, or better management of care for persons with chronic medical conditions. These efforts are complicated, however, by the fact that research has shown that being in the top of the expenditure distribution is not something that is highly persistent over time. Although expenditures in one year are correlated with expenditures in the next, there are a number of factors that determine individuals' levels of spending from one year to the next, and simply knowing base year expenditures does not mean insurers or providers can identify which individuals are most appropriate for additional attention. Nonetheless, to the extent there are specific, treatable conditions that are associated with persistently high expenditures there may be opportunities to develop methods of managing treatment that can enhance efficiency without sacrificing, or perhaps even improving, quality of care. This analysis builds on previous work on the prediction and concentration of expenditures to examine which health conditions and health status measures, in addition to demographic characteristics, are most associated with high medical care costs. The focus will be on chronic conditions, such as heart disease, cancer, diabetes, and depression, and combinations of those chronic conditions, to determine which set of conditions and other factors, including overall health status measures such as risk scores and self assessed health, are the best predictors of being in the upper tail of the expenditure distribution. For this study we pool data from the 1996 through 2004 Medical Expenditure Panel Survey (MEPS) and predict annual expenditures for individuals with single and multiple chronic conditions. The focus will be on identifying the subset of chronic conditions or combination of conditions that appear to have the greatest potential for efficiency improvement due to the level of expenditures associated with them.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".