Evaluating evidence for a neuropsychological toolkit to predict cognitive decline in PD: A systematic review
Bibliographic record
Abstract
Objectives: Neuropsychological measures used to assess cognition in Parkinson’s disease (PD) vary greatly across clinical and research settings. We conducted a systematic review to evaluate the literature pertaining to neuropsychological tools predictive of cognitive decline in PD, with a view to developing an evidence-based harmonized toolkit. Method: Following PRISMA guidelines, systematic literature searches for neuropsychological predictors of longitudinal cognitive decline in PD were performed for articles published up to August 2024 in PubMed, SCOPUS, Medline, PyscINFO and CINAHL databases. Quality was assessed using the Newcastle-Ottawa scale for individual studies and the GRADE system for each cognitive outcome. Results: Thirty-one relevant articles met inclusion criteria, with low to moderate risk of bias. Category fluency, Symbol Digit Modalities Test, Trail-making Test part A, Stroop word or color, immediate verbal memory, and Montreal Cognitive Assessment produced the highest grade of evidence (moderate), strongly supporting their predictive utility in PD. Stroop word-color, Letter Number Sequencing, pentagon copying, Trail-making Test part B, and delayed verbal and visual memory produced low quality evidence supporting their predictive utility in PD. Digit span forward and backward measures produced very low quality evidence, with consistent evidence against their predictive utility. Twelve additional measures produced very low quality of evidence due to insufficient studies or mixed results. Conclusions: The evidence base for key neuropsychological measures sensitive to cognitive decline in PD was evaluated in this systematic review. The findings will inform evidence-based tool selection for cognitive evaluations in PD and a PD-specific harmonized cognitive toolkit.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.096 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.006 | 0.002 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".