Development and internal validation of a multimorbidity index that predicts healthcare utilisation using the Canadian Longitudinal Study on Aging
Bibliographic record
Abstract
OBJECTIVES: We aimed to develop and internally validate a measure of multimorbidity burden using data from the Canadian Longitudinal Study on Aging (CLSA). DESIGN: Data from 40 264 CLSA participants (52% men) aged 45-85 years (a mean of 63 years) were analysed. We used logistic regression models to predict overnight hospitalisation in the last 12 months in the development dataset (random two-thirds of the total) and used these to construct 10 multimorbidity indices (5 models, each treated with and without an age interaction term). Thirty-five chronic conditions were considered for inclusion in these models, in addition to age and sex. We assessed predictive and convergent validity for these 10 different multimorbidity indices in the validation dataset (remaining one-third of the total). RESULTS: The absolute count of chronic conditions plus an interaction with age, displayed strong calibration properties, outperforming other candidate indices. Discrimination was modest for all of the indices that we internally validated, with C-statistics ranging from 0.66 to 0.68. The indices showed weak correlations (ie, convergent validity) with satisfaction with life, functional disability and mental health (absolute Pearson's correlation coefficients ranging from 0.11 to 0.30) but generally moderate correlations with self-rated general health (0.32-0.45). CONCLUSIONS: We investigated alternative methods to measure the multimorbidity burden of individuals, tailored to the CLSA. Our findings show that an absolute count of conditions, along with an age interaction term, has the strongest calibration for overnight hospitalisation in the last 12 months. The utility of an age interaction term in measuring multimorbidity burden may be applicable to the study of chronic disease in cohorts other than the CLSA.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.036 | 0.049 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".