Detecting agitation and aggression in persons living with dementia: a systematic review of diagnostic accuracy
Bibliographic record
Abstract
OBJECTIVE: 40-60% of persons living with dementia (PLWD) experience agitation and/or aggression symptoms. There is a need to understand the best method to detect agitation and/or aggression in PLWD. We aimed to identify agitation and/or aggression tools that are validated against a reference standard within the context of PLWD. METHODS: Our study was registered on PROSPERO (CRD42020156708). We searched MEDLINE, Embase, and PsycINFO up to April 22, 2024. There were no language or date restrictions. Studies were included if they used any tools or questionnaires for detecting either agitation or aggression compared to a reference standard among PLWD, or any studies that compared two or more agitation and/or aggression tools in the population. All screening and data extraction were done in duplicates. Study quality was assessed using the Quality Assessment of Diagnostic Accuracy Studies 2 (QUADAS-2) tool. Data extraction was completed in duplicates by two independent authors. We extracted demographic information, prevalence of agitation and/or aggression, and diagnostic accuracy measures. We also reported studies comparing the correlation between two or more agitation and/or aggression tools. RESULTS: 6961 articles were screened across databases. Six articles reporting diagnostic accuracy measures compared to a reference standard and 30 articles reporting correlation measurements between tools were included. The agitation domain of the Spanish NPI demonstrated the highest sensitivity (100%) against the agitation subsection of the Spanish CAMDEX. Single-study evidence was found for the diagnostic accuracy of commonly used agitation scales (BEHAVE-AD, NPI and CMAI). CONCLUSIONS: The agitation domain of the Spanish NPI, the NBRS, and the PAS demonstrated high sensitivities, and may be reasonable for clinical implementation. However, a limitation to this finding is that despite an extensive search, few studies with diagnostic accuracy measurements were identified. Ultimately, more research is needed to understand the diagnostic accuracy of agitation and/or aggression detection tools among PLWD.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.023 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".