Development of a method for quantifying cognitive ability in the elderly using adaptive testing
Bibliographic record
Abstract
With the impending aging of the Canadian population, the need for an assessment tool that can accurately measure cognitive ability in the elderly and monitor changes in cognition over time is rapidly gaining importance.The objective of the present study was to contribute evidence to the interpretability of scores from a novel cognitive assessment tool (GRACE) designed to be administered adaptively in the elderly population.Responses to items from two cognitive screening tests administered to patients attending a Geriatric Cognitive Disorders Clinic in Montreal were calibrated onto an interval scale using Rasch analysis.The hierarchy of items, organized by level of cognitive difficulty, was administered in a pilot adaptive format to a new cohort of patients, followed by administration of the remaining items to calculate total test scores.The reliability and validity of the GRACE method were demonstrated by comparing scores obtained from different orders of item administration (i.e.across cohorts) and by comparing scores with validated measures used in the clinic (i.e.across test methods), respectively.Additionally, demonstration of the validity of administering only a subset of items provided sound justification for developing and prospectively evaluating an optimal algorithm for adaptively administering test items.In addition to reducing test burden, the GRACE method provides a quantitative estimate of cognitive ability across the range of ability levels from normal to severe dementia and may be useful to clinicians as single tool to rapidly quantify and monitor cognitive ability.This project would not have been possible without the enthusiastic support of the clinicians and staff of the Geriatric Cognitive Disorders Clinics of the MUHC, especially their willingness to adapt their methods test administration in support of this work.I am grateful to Guylaine Bachand for inviting me to observe patient interviews and for allowing me to practice test administration under her guidance and supervision.In addition, a sincere thank you goes out to all clinicians and staff of the RVH Geriatric Day Hospital team for creating such a welcoming and supportive work environment, and for being my family-awayfrom-home over the past two years.Last but not least, special thanks go out to my family, friends, and Jesse for their love and support.This incredible experience would not have been possible without the enthusiasm, patience, and encouragement of the people who are most important in my life.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".