Identifying subgroups of adult high-cost health care users: a retrospective analysis
Bibliographic record
Abstract
<h3>Background:</h3> Few studies have categorized high-cost patients (defined by accumulated health care spending above a predetermined percentile) into distinctive groups for which potentially actionable interventions may improve outcomes and reduce costs. We sought to identify homogeneous groups within the persistently high-cost population to develop a taxonomy of subgroups that may be targetable with specific interventions. <h3>Methods:</h3> We conducted a retrospective analysis in which we identified adults (≥ 18 yr) who lived in Alberta between April 2014 and March 2019. We defined “persistently high-cost users” as those in the top 1% of health care spending across 4 data sources (the Discharge Abstract Database for inpatient encounters; Practitioner Claims for outpatient primary care and specialist encounters; the Ambulatory Care Classification System for emergency department encounters; and the Pharmaceutical Information Network for medication use) in at least 2 consecutive fiscal years. We used latent class analysis and expert clinical opinion in tandem to separate the persistently high-cost population into subgroups that may be targeted by specific interventions based on their distinctive clinical profiles and the drivers of their health system use and costs. <h3>Results:</h3> Of the 3 919 388 adults who lived in Alberta for at least 2 consecutive fiscal years during the study period, 21 115 (0.5%) were persistently high-cost users. We identified 9 subgroups in this population: people with cardiovascular disease (<i>n</i> = 4537; 21.5%); people receiving rehabilitation after surgery or recovering from complications of surgery (<i>n</i> = 3380; 16.0%); people with severe mental health conditions (<i>n</i> = 3060; 14.5%); people with advanced chronic kidney disease (<i>n</i> = 2689; 12.7%); people receiving biologic therapies for autoimmune conditions (<i>n</i> = 2538; 12.0%); people with dementia and awaiting community placement (<i>n</i> = 2520; 11.9%); people with chronic obstructive pulmonary disease or other respiratory conditions (<i>n</i> = 984; 4.7%); people receiving treatment for cancer (<i>n</i> = 832; 3.9%); and people with unstable housing situations or substance use disorders (<i>n</i> = 575; 2.7%). <h3>Interpretation:</h3> Using latent class analysis supplemented with expert clinical review, we identified 9 policy-relevant subgroups among persistently high-cost health care users. This taxonomy may be used to inform policy, including identifying interventions that are most likely to improve care and reduce cost for each subgroup.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".