Limitations for Interpreting Failure on Individual Subtests of the Montreal Cognitive Assessment
Bibliographic record
Abstract
BACKGROUND: The Montreal Cognitive Assessment (MoCA) is sensitive to mild forms of cognitive impairment in geriatric populations and asks questions under the subheadings visuospatial/executive, naming, attention, language, abstraction, delayed recall, and orientation. This study examined the extent to which these subsets of MoCA items evaluate their intended cognitive domains. METHODS: Clinical data from 185 geriatric memory clinic outpatients who underwent cognitive screening and subsequent neuropsychological assessment were analyzed. Factor analysis of their neuropsychological test scores identified 5 cognitive domains memory, language, visuospatial ability, attention/processing speed, and cognitive control. Scores on MoCA subtests were examined for their correlations with individual factor scores and for their sensitivity and specificity in predicting impairment within each domain. RESULTS: The MoCA subtest scores correlated significantly but modestly with neuropsychological test factor scores in their corresponding domains, for example, the correlation between 5-word recall and the memory factor was 0.46. However, subtest scores were poor predictors of impaired performance on the tests contributing to each cognitive domain. The best predictive accuracy was seen for the visuospatial/executive subtest that showed fair accuracy at predicting impairment on tests in the visuospatial domain. Other subtests showed unacceptably poor levels of accuracy when predicting impaired scores in their respective domains (60%-67%). CONCLUSIONS: In a sample of geriatric outpatients referred for cognitive assessment, performance on individual items and subtests of the MoCA yields insufficient information to draw conclusions about impairment in specific cognitive domains as determined by neuropsychological testing.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".