Handwriting in Mild Cognitive Impairment: Reliability Assessment and Machine Learning–Based Screening
Bibliographic record
Abstract
Background Mild cognitive impairment (MCI) is a precursor of dementia. Therefore, MCI identification and monitoring are crucial to delaying dementia onset. Given the limits of existing clinical tests, objective support tools are needed. Objective This work investigates quantitative handwriting analysis, tailored to enable domestic monitoring, as a noninvasive approach for MCI screening and assessment. Methods A sensorized ink pen, used on paper and equipped with sensors, memory, and a communication unit, was used for data acquisition. The tasks included writing a grocery list and free text to mimic daily life handwriting, and a clinical dictation test (parole-non-parole [PnP] test), featuring regular, irregular, and made-up words, aimed at assessing MCI dysgraphia. From the recorded data, 106 indicators describing the performance in terms of time, fluency, exerted force, and pen inclination were computed. A total of 57 patients with MCI were recruited, of whom 45 performed a test-retest protocol. The indicators were examined to assess their test-retest reliability. The indicators from the test repetition were used to assess their relationship with the scores of clinical tests via correlation analysis. For the PnP test, differences in the indicators among the 3 types of words were statistically investigated. These analyses were conducted separately for the cursive (2/3 of the sample) and block letters (1/3 of the sample) allographs, with the level of significance set at 5%. Data from healthy older adults were available for the grocery list (34 participants) and free text (45 participants) tasks. These were exploited to build machine learning classification models for the distinction between patients with MCI and healthy controls. Results When dealing with reliability, 93% and 44% of the indicators were characterized by a significant reliability of at least moderate intensity for cursive and block letters respectively. As for the correlation analysis, patients with preserved cognitive status and daily life functionality were associated with significantly better temporal performances, both in free writing and PnP. The analysis of PnP highlighted the presence of surface dysgraphia in the recruited sample, as irregular words showed significantly worse temporal indicators with respect to regular and made-up ones. The classification models’ built-in free writing data achieved accuracies ranging from 0.80 to 0.93 and F1-scores from 0.81 to 0.92 according to the input dataset. Conclusions The presented results suggest the suitability of ecological handwriting analysis for the all-around monitoring of MCI, from early screening to disease progression evaluation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".