Accuracy of Observer-Rated Measurement Scales for Depression Assessment in Patients with Major Neurocognitive Disorders Residing in Long-Term Care Centers: A Systematic Review
Bibliographic record
Abstract
INTRODUCTION: Depression is often under-detected in long-term care (LTC) patients with major neurocognitive disorders (MNCD) and is associated with important morbidity, mortality, and costs. Observer-rated outcome measures (ObsROMs) could help resolve this problematic; however, evidence on their accuracy is scattered in the literature. This systematic review aimed at summarizing this evidence. METHODS: A literature search was conducted in 7 databases using keywords, MeSHs, and bibliographic searches. We included studies published before January 2022 and reporting on the accuracy of a depression ObsROM used in LTC patients with MNCD. Data extraction, analysis, synthesis, and study methodological quality assessments were done by two authors, and discrepancies were resolved by consensus. RESULTS: Among 9,660 articles retrieved, 8 studies reporting on 11 depression measures were included. Scales were classified as patient-reported outcome measures used as Obs-ROMs or true ObsROMs. Among the first category, the Cornell Scale for Depression in Dementia (CSDD) and the Montgomery-Asberg Depression Rating Scale (MADRS) performed best (area under the curve [AUC]: 0.73-0.87), although both presented with low positive predictive values and high negative predictive values. Among the second category, the Nursing Homes Short Depression Inventory (NH-SDI) performed best, with an AUC of 0.93 and ≥85% sensitivity, specificity, and predictive values. CONCLUSION: The CSDD and MADRS may be useful to rule out depression in LTC patients with MNCD, whereas the NH-SDI may be useful to rule in and out depression within this same population. Before recommending their use, adequately powered studies to further examine their accuracy in different contexts are necessary.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".