Prospective comparison of three methods for detecting peri‐operative neurocognitive disorders in older adults undergoing cardiac and non‐cardiac surgery
Bibliographic record
Abstract
Postoperative neurocognitive disorders occur frequently in older adult patients. Neuropsychological assessment is the gold standard for diagnosis, but the resources required for routine use are significant. Instead, it is common for simplified and unvalidated tests to be used for trials and in clinical practice. We undertook a single-centre prospective observational study in elective surgical patients aged ≥ 65 years recruited between September 2019 and January 2021. Patients underwent neuropsychological assessment, the Modified Telephone Interview for Cognitive Status and Montreal Cognitive Assessment before surgery. Tests were repeated at approximately four to eight postoperative weeks. We included 105 patients and 28 (27%) were lost to follow-up. Pre-operative Modified Telephone Interview for Cognitive Status and cognitive domain scores were very weakly to moderately correlated (r = 0.09-0.41). Pre-operative Montreal Cognitive Assessment and cognitive domain scores were very weakly to weakly correlated (r = 0.17-0.37) Postoperative Modified Telephone Interview for Cognitive Status and cognitive domain scores were very weakly to weakly correlated (r = 0.09-0.36). Postoperative Montreal Cognitive Assessment score and cognitive domain scores were very weakly to weakly correlated (r = 0.07-0.36). Overall, there was limited agreement between tests. We conclude that the Modified Telephone Interview for Cognitive Status and Montreal Cognitive Assessment should not be used in isolation to diagnose postoperative neurocognitive disorders. There seems to be little to no pre-operative, postoperative or pre- to postoperative correlation between these tests and the neuropsychological assessment in older adults without pre-operative cognitive impairment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.007 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".