Development and validation of a score to assess complexity of general internal medicine patients at hospital discharge: a prospective cohort study
Bibliographic record
Abstract
OBJECTIVE: We aimed to develop and validate a score to assess inpatient complexity and compare its performance with two currently used but not validated tools to estimate complexity (ie, Charlson Comorbidity Index (CCI), patient clinical complexity level (PCCL)). METHODS: Consecutive patients discharged from the department of medicine of a tertiary care hospital were prospectively included into a derivation cohort from 1 October 2016 to 16 February 2017 (n=1407), and a temporal validation cohort from 17 February 2017 to 31 March 2017 (n=482). The physician in charge assessed complexity. Potential predictors comprised 52 parameters from the electronic health record such as health factors and hospital care usage. We fit a logistic regression model with backward selection to develop a prediction model and derive a score. We assessed and compared performance of model and score in internal and external validation using measures of discrimination and calibration. RESULTS: Overall, 447 of 1407 patients (32%) in the derivation cohort, and 116 of 482 patients (24%) in the validation cohort were identified as complex. Eleven variables independently associated with complexity were included in the score. Using a cut-off of ≥24 score points to define high-risk patients, specificity was 81% and sensitivity 57% in the validation cohort. The score's area under the receiver operating characteristic (AUROC) curve was 0.78 in both the derivation and validation cohort. In comparison, the CCI had an AUROC between 0.58 and 0.61, and the PCCL between 0.64 and 0.69, respectively. CONCLUSIONS: We derived and internally and externally validated a score that reflects patient complexity in the hospital setting, performed better than other tools and could help monitoring complex patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.020 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".