Remote Clinical Neuropsychological Assessment Using Portable Automated Rapid Testing (PART): Validation in a Healthy Older Adult Population (Preprint)
Bibliographic record
Abstract
Background Remote, scalable cognitive assessment could improve detection and longitudinal monitoring of age-related cognitive change, but few studies have validated comprehensive digital neuropsychological batteries administered entirely at home. Objective This study aims to evaluate the feasibility, acceptability, and short-term test-retest reliability of a remotely delivered digital neuropsychological battery (portable automated rapid testing [PART]) in cognitively healthy older adults. Methods We screened 82 English-speaking, cognitively healthy older adults aged 50 to 85 years and mailed configured tablets to participants. Researchers remotely administered a battery of neuropsychological assessments spanning language fluency (Boston Naming Test, verbal fluency: letter “F,” supermarket items, and animals), memory (verbal paired associates, word list recall, and logical memory recall), praxis memory (clock drawing and constructional praxis [CP]), and executive functioning (trail making test [TMT]) using the PART app twice, approximately 1 month apart. Task comfortability was summarized with a comfort index based on participants’ self-reports after completing each task. Reliability analyses included paired t tests to detect group-level shifts in performance, Pearson correlations for short-term stability, and Bland-Altman limits of agreement (LoA) to quantify bias and within-person variability. Performance across all tasks was also compared to comparative samples from large normative studies using equivalent paper-and-pencil versions of these tasks. Results A total of 72 participants completed the first time point (T1), 64 completed the second time point (T2), and the analytic sample comprised a total of 63 participants. Overall comfort was high across sessions (mean comfort index was 88% at T1 and 84% at T2). Memory and executive function measures showed moderate test-retest correlations (r=0.39-0.73), while clock drawing and CP showed low or nonsignificant associations, likely due to ceiling effects. Paired tests indicated small but significant practice effects for verbal paired associates (t62=−3.10; P=.003) and word list (t57=−2.50; P=.02), with substantial practice effects for logical memory (story A: t56=−2.95; P=.005) and TMT part A (t50=−3.97; P<.001) and part B (t42=−2.05; P=.05). Bland-Altman analyses revealed minimal mean bias for many tasks but notably wide LoA for TMT (part A: LoA=−44.7 to 29.2 seconds; part B: LoA=−78.1 to 57.9 seconds). Importantly, in descriptive comparisons, our sample generally overlapped within ±1 SD of the medians in comparative paper-and-pencil normative samples. Conclusions PART was well tolerated among healthy older adults and reproduces group-level normative patterns with moderate short-term reliability for several conventional measures. However, nontrivial within-person variability, practice effects, ceiling effects, and wide LoAs for some measures (eg, TMT, CP, and clock drawing) limit interpretation of individual change. Future work should validate alternative digital outcome metrics (eg, stylus interactions and speech recordings), extend sampling to more diverse and lower-education cohorts, and evaluate more ecologically valid designs to improve sensitivity for detecting meaningful within-person cognitive fluctuations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | no category Domain: not available · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Observational | low |
| gpt | no category Domain: not available · Genre: Methods About the Canadian research system: no · About a Canadian topic: no | Observational | low |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".