Evaluating evidence for a neuropsychological toolkit to predict cognitive decline in PD: A systematic review
Bibliographic record
Abstract
Objectives: Neuropsychological measures used to assess cognition in Parkinson’s disease (PD) vary greatly across clinical and research settings. We conducted a systematic review to evaluate the literature pertaining to neuropsychological tools predictive of cognitive decline in PD, with a view to developing an evidence-based harmonized toolkit. Method: Following PRISMA guidelines, systematic literature searches for neuropsychological predictors of longitudinal cognitive decline in PD were performed for articles published up to August 2024 in PubMed, SCOPUS, Medline, PyscINFO and CINAHL databases. Quality was assessed using the Newcastle-Ottawa scale for individual studies and the GRADE system for each cognitive outcome. Results: Thirty-one relevant articles met inclusion criteria, with low to moderate risk of bias. Category fluency, Symbol Digit Modalities Test, Trail-making Test part A, Stroop word or color, immediate verbal memory, and Montreal Cognitive Assessment produced the highest grade of evidence (moderate), strongly supporting their predictive utility in PD. Stroop word-color, Letter Number Sequencing, pentagon copying, Trail-making Test part B, and delayed verbal and visual memory produced low quality evidence supporting their predictive utility in PD. Digit span forward and backward measures produced very low quality evidence, with consistent evidence against their predictive utility. Twelve additional measures produced very low quality of evidence due to insufficient studies or mixed results. Conclusions: The evidence base for key neuropsychological measures sensitive to cognitive decline in PD was evaluated in this systematic review. The findings will inform evidence-based tool selection for cognitive evaluations in PD and a PD-specific harmonized cognitive toolkit.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.035 | 0.136 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.013 | 0.014 |
| Bibliometrics | 0.022 | 0.016 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.006 | 0.005 |
| Open science | 0.004 | 0.003 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".