Accuracy of the short‐form Montreal Cognitive Assessment: Systematic review and validation
Bibliographic record
Abstract
INTRODUCTION: Short-form versions of the Montreal Cognitive Assessment (SF-MoCA) are increasingly used to screen for dementia in research and practice. We sought to collate evidence on the accuracy of SF-MoCAs and to externally validate these assessment tools. METHODS: We performed systematic literature searching across multidisciplinary electronic literature databases, collating information on the content and accuracy of all published SF-MoCAs. We then validated all the SF-MoCAs against clinical diagnosis using independent stroke (n = 787) and memory clinic (n = 410) data sets. RESULTS: We identified 13 different SF-MoCAs (21 studies, n = 6477 participants) with differing test content and properties. There was a pattern of high sensitivity across the range of SF-MoCA tests. In the published literature, for detection of post stroke cognitive impairment, median sensitivity across included studies: 0.88 (range: 0.70-1.00); specificity: 0.70 (0.39-0.92). In our independent validation using stroke data, median sensitivity: 0.99 (0.80-1.00); specificity: 0.40 (0.14-0.87). To detect dementia in older adults, median sensitivity: 0.88 (0.62-0.98); median specificity: 0.87 (0.07-0.98) in the literature and median sensitivity: 0.96 (range: 0.72-1.00); median specificity: 0.36 (0.14-0.86) in our validation. Horton's SF-MoCA (delayed recall, serial subtraction, and orientation) had the most favorable properties in stroke (sensitivity: 0.90, specificity: 0.87, positive predictive value [PPV]: 0.55, and negative predictive value [NPV]: 0.93), whereas Cecato's "MoCA reduced" (clock draw, animal naming, delayed recall, and orientation) performed better in the memory clinic (sensitivity: 0.72, specificity: 0.86, PPV: 0.55, and NPV: 0.93). CONCLUSIONS: There are many published SF-MoCAs. Clinicians and researchers using a SF-MoCA should be explicit about the content. For all SF-MoCA, sensitivity is high and similar to the full scale suggesting potential utility as an initial cognitive screening tool. However, choice of SF-MoCA should be informed by the clinical population to be studied.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.147 | 0.456 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.010 | 0.012 |
| Bibliometrics | 0.016 | 0.014 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.006 | 0.005 |
| Open science | 0.004 | 0.003 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".