Best Definitions of Multimorbidity to Identify Patients With High Health Care Resource Utilization
Bibliographic record
Abstract
ObjectiveTo compare different definitions of multimorbidity to identify patients with higher health care resource utilization.Patients and MethodsWe used a multinational retrospective cohort including 147,806 medical inpatients discharged from 11 hospitals in 3 countries (United States, Switzerland, and Israel) between January 1, 2010, and December 31, 2011. We compared the area under the receiver operating characteristic curve (AUC) of 8 definitions of multimorbidity, based on International Classification of Diseases codes defining health conditions, the Deyo-Charlson Comorbidity Index, the Elixhauser-van Walraven Comorbidity Index, body systems, or Clinical Classification Software categories to predict 30-day hospital readmission and/or prolonged length of stay (longer than or equal to the country-specific upper quartile). We used a lower (yielding sensitivity ≥90%) and an upper (yielding specificity ≥60%) cutoff to create risk categories.ResultsDefinitions had poor to fair discriminatory power in the derivation (AUC, 0.61-0.65) and validation cohorts (AUC, 0.64-0.71). The definitions with the highest AUC were number of (1) health conditions with involvement of 2 or more body systems, (2) body systems, (3) Clinical Classification Software categories, and (4) health conditions. At the upper cutoff, sensitivity and specificity were 65% to 79% and 50% to 53%, respectively, in the validation cohort; of the 147,806 patients, 5% to 12% (7474 to 18,008) were classified at low risk, 38% to 55% (54,484 to 81,540) at intermediate risk, and 32% to 50% (47,331 to 72,435) at high risk.ConclusionOf the 8 definitions of multimorbidity, 4 had comparable discriminatory power to identify patients with higher health care resource utilization. Of these 4, the number of health conditions may represent the easiest definition to apply in clinical routine. The cutoff chosen, favoring sensitivity or specificity, should be determined depending on the aim of the definition.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".