Delirium detection in the emergency department: A diagnostic accuracy meta‐analysis of history, physical examination, laboratory tests, and screening instruments
Bibliographic record
Abstract
INTRODUCTION: Geriatric emergency department (ED) guidelines emphasize timely identification of delirium. This article updates previous diagnostic accuracy systematic reviews of history, physical examination, laboratory testing, and ED screening instruments for the diagnosis of delirium as well as test-treatment thresholds for ED delirium screening. METHODS: We conducted a systematic review to quantify the diagnostic accuracy of approaches to identify delirium. Studies were included if they described adults aged 60 or older evaluated in the ED setting with an index test for delirium compared with an acceptable criterion standard for delirium. Data were extracted and studies were reviewed for risk of bias. When appropriate, we conducted a meta-analysis and estimated delirium screening thresholds. RESULTS: Full-text review was performed on 55 studies and 27 were included in the current analysis. No studies were identified exploring the accuracy of findings on history or laboratory analysis. While two studies reported clinicians accurately rule in delirium, clinician gestalt is inadequate to rule out delirium. We report meta-analysis on three studies that quantified the accuracy of the 4 A's Test (4AT) to rule in (pooled positive likelihood ratio [LR+] 7.5, 95% confidence interval [CI] 2.7-20.7) and rule out (pooled negative likelihood ratio [LR-] 0.18, 95% CI 0.09-0.34) delirium. We also conducted meta-analysis of two studies that quantified the accuracy of the Abbreviated Mental Test-4 (AMT-4) and found that the pooled LR+ (4.3, 95% CI 2.4-7.8) was lower than that observed for the 4AT, but the pooled LR- (0.22, 95% CI 0.05-1) was similar. Based on one study the Confusion Assessment Method for the Intensive Care Unit (CAM-ICU) is the superior instrument to rule in delirium. The calculated test threshold is 2% and the treatment threshold is 11%. CONCLUSIONS: The quantitative accuracy of history and physical examination to identify ED delirium is virtually unexplored. The 4AT has the largest quantity of ED-based research. Other screening instruments may more accurately rule in or rule out delirium. If the goal is to rule in delirium then the CAM-ICU or brief CAM or modified CAM for the ED are superior instruments, although the accuracy of these screening tools are based on single-center studies. To rule out delirium, the Delirium Triage Screen is superior based on one single-center study.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.026 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.003 | 0.001 |
| Bibliometrics | 0.002 | 0.004 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".