Validity and Reliability of Outcome Measurement Instruments for Cognitive Function in Alzheimer’s Disease: A Systematic Review
Bibliographic record
Abstract
INTRODUCTION: In this systematic review, we aimed to identify suitable assessment and measurement tools for screening individuals for cognitive impairment and Alzheimer's disease (AD). We conducted a comprehensive evaluation of the reliability and validity of cognitive function assessment instruments. Based on our findings, we offer insightful suggestions for further research on cognitive function scale development and clinical researchers in AD. METHODS: We searched the PubMed and CNKI databases for community-based studies aimed at developing or evaluating the validity or reliability of cognitive function assessment scales. Only studies written in English and Chinese that reported the development of cognitive function assessment scales and/or the validation of cognitive impairment severity in patients with cognitive impairment and AD were eligible for inclusion. The methodological quality of the studies, based on reliability (i.e., internal consistency and test-retest reliability) and validity (i.e., construct validity), was assessed using the consensus-based standards for the selection of health measurement instruments (COSMIN) according to the "worst score counts" principle. Subsequently, the measurement properties were rated qualitatively. Results were summarized and rated using the modified Grading of Recommendations, Assessment, Development, and Evaluation. Based on the results, recommendations were categorized into four levels: A, B, C, and D. RESULTS: We retrieved a total of 804 studies. Following screening, a total of 62 articles were included, which reported 49 cognitive impairment assessment scales. The methodological quality of studies ranged from inadequate to very good, and the measurement properties varied from sufficient (+) to indeterminate (?). We found that the AD Assessment Scale-cognitive subscale (ADAS-Cog), Montreal Cognitive Assessment (MoCA), Baylor Profound Mental Status Examination (BPMSE), Clinical Dementia Rating (CDR), and the other 28 scales had sufficient validity and reliability. CONCLUSION: Our evaluation according to the COSMIN guidelines suggested that the ADAS-Cog, MoCA, BPMSE, CDR scale, and Mini-mental State Examination could be used to assess the degree of cognitive impairment in patients with AD. When developing cognitive function assessment scales, factors such as time and linguistic and cultural differences could be carefully considered.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".