MétaCan
Menu
Back to cohort
Record W7117716984 · doi:10.17605/osf.io/qan4m

Screening for Post-stroke Cognitive Impairment: a comparison of univariate and multivariate normative approaches

2025· other· W7117716984 on OpenAlexaboutno aff
Céline Gillebert, Hanne Huygelier, Nele Demeyere, Sam Sappho Webb

Bibliographic record

VenueOpen Science Framework · 2025
Typeother
Language
Field
Topic
Canadian institutionsnot available
Fundersnot available
KeywordsNormativeCognitionCognitive impairmentUnivariateStroke (engine)AnosognosiaCognitive Assessment SystemNeuropsychologyLogistic regression

Abstract

fetched live from OpenAlex

Screening for domain-specific cognitive impairment after stroke is essential due to the heterogeneous patterns and presentation of cognitive impairments post-stroke and the effect of these impairments on daily life (Esmael et al., 2021; Jokinen et al., 2015; Kusec et al., 2023; Merriman et al., 2019; Milosevich et al., 2023; Mole & Demeyere, 2020; Nys et al., 2005, 2007; Oksala et al., 2009; Samuelsson et al., 2021; Sexton et al., 2019a; Stroke Association, 2018). Failure to detect impairments can lead to failure to target rehabilitation to address or attenuate these problems, and can increase the severity of disability and incidence of mortality (Lucka et al., 2022). To screen for post-stroke cognitive impairment (PSCI) many approaches are in use (Sexton et al., 2019b). Some researchers identify a subtest-specific “impairment” by comparing the patient’s score to a cut-off based on normative data (e.g., the 5th centile cut off, or a z-score > 1 SD from the mean) (e.g., (Demeyere et al., 2015; Sexton et al., 2019b), and then identify all patients with PSCI as those who show impairment on at least one test (e.g., Demeyere et al., 2016; Jokinen et al., 2015; Nys et al., 2007). An alternative approach is to combine subtest scores in one global score indicating PSCI, such as summing subtests of cognitive screens such as the Montreal Cognitive Assessment (MoCA) (Nasreddine et al., 2005). When identifying PSCI as “at least 1 impairment”, one may risk inflating false positives as the approach typically does not control for multiple comparisons. Indeed, in the seminal paper by Brooks and colleagues (Brooks et al., 2009), they highlight that, even in healthy adults, there is an expected high frequency of low scores when completing many tests. This frequency will increase if the threshold for impairments is liberal (set such that many are going to be impaired). Consequently, the number of individuals with at least 1 low score increases the more tests there are administered. In contrast, when identifying PSCI using total scores, one risks under-identifying domain-specific impairments (Demeyere et al., 2016). A second challenge is the fact that a patient may score subclinically on multiple tests, not meeting the per-test criterion of impairment, while their overall test profile may still deviate from healthy controls, and thus with current per-test criteria (i.e., univariate testing) lead to under identification of impairment (Huizenga et al., 2007). Thus, assessors face the challenge of assessing cognition sufficiently to identify potential domain-specific impairments, while: 1) avoiding false positives, and 2) avoiding false negatives resulting from univariate testing. Since the early 2000s, more attention has been given to offer solutions for this challenge. For instance, more neuropsychological tests are considering a multivariate base rate of performance in normative data (Kiselica et al., 2024) and examples of these base rates to aid interpretation of performance have been developed, including for the Wechsler Adult Intelligence Scale–Fourth Edition (Brooks et al., 2013). The multivariate base rates allow to control for false positives, while considering test associations in normative samples. In addition, a test statistic, the Multivariate Normative Comparison (MNC), has been validated to compare a patient’s test profile to normative data in a multivariate way. A detailed tutorial has been written regarding how to compute MNC for psychological test data in order to understand patterns of performance rather than just isolated subtest performance (Huizenga et al., 2007). Huizenga et al demonstrated that in situations where the number of (sub)tests is just exceeded by the number of controls the Bonferroni correction for multiple comparisons of tests is sufficient to see if a given patient is different to normative performance across multiple tests (Huizenga et al., 2007). Whereas, in situations with more healthy control data, where estimations can be more precise and patterns of performance across different (sub)tests can be estimated, MNC is more statistically powerful to identify cognitive impairment (Huizenga et al., 2007). Although the MNC approach may be theoretically more powerful to identify PSCI across multiple domains, it has not yet been empirically evaluated for identifying PSCI. It is thus essential to investigate whether multivariate methods of interpreting cognitive tests impact the identification of PSCI. To this end, we will investigate different methods of identifying across-domain PSCI using the Oxford Cognitive Screen (OCS). The OCS (Demeyere et al., 2015) is widely used in clinical practice as a first line screen for cognitive impairment (Murphy et al., 2023). The OCS briefly screens for impairments in language, memory, attention, executive function, praxis, and numerical cognition. At its base level, the OCS only provides impairment scores on subtests and does not officially have cut-offs for across-domain PSCI. In the current study, we will contrast four methods to identify PSCI: (1) the traditional approach of identifying PSCI based on an “at least 1 impairment” criterion and 5th centile cut-offs per test (not corrected for multiple comparisons) (Demeyere et al., 2016; Kusec et al., 2023; Sexton et al., 2019b), (2) an adjusted version of the “at least 1 impairment” approach where the per-test cut-offs were Bonferroni corrected, (3) a total score approach which combines all subtests in a single outcome measure and contrasts this to a normative group, and (4) the “MNC” approach which has the power to consider test-associations in the normative sample and identify multivariate deviations in test performance (Grasman et al., 2010). Our primary aim was to assess whether these methods result in different estimates of the prevalence of PSCI in a large stroke sample. In addition, we examined for how many patients’ diagnosis would be different depending on the approach used. To estimate the impact of multiple comparisons for the OCS, we applied these methods on the normative group as well. We predict that the “at least 1 impairment” approach without Bonferroni corrected cut-offs for subtest impairment will not control the false positive rate (i.e., maintain PSCI identification at the nominal significance level in the normative group). In a worst-case (unrealistic) scenario, PSCI would be identified in 45% of the normative group (if the 13 OCS subtests were fully independent) (Huizenga et al., 2007). We predict that all other methods will control this rate at the nominal significance level. For the stroke patients, we predict that the MNC approach will identify PSCI in more patients than the Bonferroni corrected “at least 1 impairment” approach (as it allows to consider test-associations) and in less patients than the uncorrected “at least 1 impairment” approach (as it controls for multiple comparisons).

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.009
metaresearch head score (Gemma)0.010
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Meta-epidemiology (narrow), Science and technology studies, Scholarly communication, Open science, Research integrity, Insufficient payload (model declined to judge)
Consensus categoriesMeta-epidemiology (narrow), Science and technology studies, Open science
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Qualitative · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.592
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0090.010
Meta-epidemiology (narrow)0.0020.002
Meta-epidemiology (broad)0.0040.000
Bibliometrics0.0030.006
Science and technology studies0.0030.009
Scholarly communication0.0030.003
Open science0.0080.009
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.086
GPT teacher head0.383
Teacher spread0.297 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designQualitative
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueOpen Science FrameworkFrench-language works237,207