Screening for older inpatients at risk for long length of stay: which clinical tool to use?
Bibliographic record
Abstract
BACKGROUND: Screening for inpatients at risk for long length of stay (LOS) is the first step of an effective hospital care plan for older inpatients. This study aims, in older adults admitted to a geriatric acute care ward, to examine and compare the 6-item brief geriatric assessment (BGA) and the "Programme de Recherche sur l'Intégration des Services pour le Maintien de l'Autonomie" (PRISMA-7) risk levels with long LOS, and to establish their performance criteria (i.e., sensitivity, specificity, positive predictive value, negative predictive value, likelihood ratios) for LOS. METHODS: Based on an observational, retrospective, cohort design, 166 inpatients aged ≥75 admitted to a geriatric acute care ward of a McGill University-affiliated hospital (Montreal, Quebec, Canada) were recruited. The risk levels of the 6-item BGA (low, moderate and high) and the PRISMA-7 (low versus high) were calculated from a baseline assessment. The LOS was subsequently calculated in number of days. RESULTS: Only the 6-item BGA high risk level was associated with a long LOS (Odds ratio = 1.1 with P = 0.028 and Hazard ratio = 2.1 with P = 0.004). Kaplan-Meier distributions showed that there was no significant difference in the delay of hospital discharge between the low and high-risk level reported by the PRISMA-7 (P = 0.381), whereas the 6-item BGA three risk levels differed significantly (P = 0.008), with individuals at high risk levels being discharged later when compared to those with low (P = 0.001) and moderate (P = 0.019) risk levels. Both tools' performance criteria were poor (i.e., < 0.70), except for PRISMA-7's sensitivity which was 100%. CONCLUSION: The 6-item BGA risk levels were associated with LOS, low risk-level being associated with short LOS and high-risk level with long LOS, but no association was reported with the PRISMA-7 risk levels. Both tools had poor performance criteria for long LOS, suggesting that they cannot be used as prognostic tools with current scientific knowledge.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".