Comparison between adjusted Montreal Cognitive Assessment and neuropsychological assessment for diagnosing postoperative neurocognitive disorders
Bibliographic record
Abstract
The current gold standard neuropsychological assessment for detecting postoperative neurocognitive disorders is too time-consuming, costly and burdensome to use in clinical practice. Brief screening instruments, such as the Montreal Cognitive Assessment (MoCA), are used frequently instead. However, previous research by our team suggested that the original MoCA is not suitable to detect postoperative neurocognitive disorders in older adult surgical patients [1]. To improve the accuracy of the MoCA, Kessels et al. presented norms controlling for age, sex and educational level [2]. Accordingly, our study aimed to compare the performance of the adjusted MoCA score in diagnosing postoperative neurocognitive disorder. We prospectively enrolled patients aged ≥ 65 y scheduled for elective surgery, involving any type of anaesthesia or surgical procedure, from September 2019 to January 2021, after approval by our local research ethics committee. Patients who were not fluent in Dutch, had pre-operative cognitive impairment, severe hearing impairment or needed several procedures under anaesthesia were not studied. The original study is described in full elsewhere [1]. Simultaneous administration of neuropsychological assessment and MoCA occurred pre-operatively and 30–60 days postoperatively, using alternate versions to minimise practice effect. Performance on neuropsychological assessment was reported as T-scores after comparison to a Dutch norm group (https://andi.nl). For neuropsychological assessment, a decline of 1–2 SD on ≥ 1 cognitive domain score indicated mild postoperative cognitive disorder, and ≥ 2 SD decline indicated major postoperative neurocognitive disorder [3]. In the post hoc analysis, we transformed the original, education-uncorrected, MoCA scores to percentiles according to Kessels et al. [2]. Mild postoperative neurocognitive disorder was defined as a reliable change index decrease of 1–2 SD [4] and ≥ 2 SD decline indicated major postoperative neurocognitive disorder. Test–retest reliability was measured by intraclass correlation coefficient. Data were missing completely at random and were imputed. Sensitivity, specificity and area under the receiver operating characteristic curve of the adjusted MoCA were calculated. We examined pre-operative, postoperative and pre- to postoperative correlations of MoCA and total neuropsychological assessment and domain scores. We transformed the outcome to z-scores to assess agreement between MoCA and neuropsychological assessment by Bland–Altman plots. Ordinary or regression limits of agreement were chosen based on the presence or absence of proportional bias [5]. A total of 73 patients completed neuropsychological assessment and MoCA. Baseline characteristics are detailed in online Supporting Information Appendix S1. Neuropsychological assessment identified 14 (19%) cases of postoperative neurocognitive disorder and MoCA diagnosed 15 (21%) patients with cognitive disorders. Only two cases were diagnosed by both instruments (Table 1). Neuropsychological assessment classified all patients with mild postoperative neurocognitive disorder and MoCA diagnosed three patients with major cognitive disorder; however, only one of these cases was also diagnosed with postoperative neurocognitive disorder by neuropsychological assessment. Test–retest reliability of the adjusted MoCA was moderate (online Supporting Information Appendix S2). Sensitivity and specificity of the adjusted MoCA were 0.14 (95%CI 0.03–0.38) and 0.78 (95%CI 0.66–0.87), respectively. The area under the receiver operating characteristic curve was 0.54 (95%CI 0.38–0.70). The correlations between pre-operative adjusted MoCA and neuropsychological assessment domain scores were weak to moderate (r = 0.12–0.48). Postoperative correlations were very weak to weak (r = -0.03–0.28) and pre- to postoperative MoCA correlations very weak (r = -0.10–0.09) (online Supporting Information Appendix S3). There was little agreement between pre-operative and postoperative MoCA scores compared with total neuropsychological assessment scores as well as domain scores (Fig. 1, online Supporting Information Appendices S4 and S5). Our results suggest that the MoCA, despite adjustments for age, sex and educational level, is inadequate for diagnosing postoperative neurocognitive disorders in older adult elective surgical patients. It should not be used for clinical or research purposes for postoperative neurocognitive disorders, aligning with our previous research [1]. Sensitivity and specificity were comparable between adjusted (0.14–0.78) and original MoCA (0.21–0.84), respectively. Possible inadequacy of the MoCA could arise because of the subtlety in cognitive change in patients with postoperative neurocognitive disorders, as MoCA is only tailored for monitoring large cognitive changes in patients with dementia [6]. Additionally, studies showed a limited correlation between MoCA items and corresponding neuropsychological assessment scores, questioning the validity of the MoCA items and their comparability with neuropsychological assessment [7]. A limitation is the lack of a uniform definition for postoperative neurocognitive dysfunction. We chose the recommended approach using cognitive domain scores, but a different definition could possibly alter the results [3]. However, we compared the two diagnosing tools without the need for a definition by measuring agreement and correlations. Furthermore, various tests are used across studies for the gold standard neuropsychological assessment [8]. A strength was that MoCA was administered by trained staff. We hypothesise that these findings extend to other brief cognitive tests, like the Mini-Mental State Exam, and, therefore, recommend caution in their use for diagnosing postoperative neurocognitive disorders. Collectively, our findings underscore the need for an adequate brief diagnostic tool tailored for postoperative neurocognitive disorder as existing brief instruments, such as the (adjusted) MoCA, seem inadequate. This study was supported by the Amsterdam University Fund. Data are available upon reasonable request from the corresponding author. No competing interests declared. Appendix S1. Baseline characteristics of included patients. Appendix S2. Intraclass correlation coefficients. Appendix S3. Correlations between adjusted Montreal Cognitive Assessment and neuropsychological assessment. Appendix S4. Bland-Altman analysis. Appendix S5. Bland-Altman plots. Please note: The publisher is not responsible for the content or functionality of any supporting information supplied by the authors. Any queries (other than missing content) should be directed to the corresponding author for the article.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".