Evaluation of Interpretable Speech Biomarkers for Monitoring Alzheimer’s Disease and Mild Cognitive Impairment Progression
Bibliographic record
Abstract
Abstract Background Alzheimer’s disease (AD) is a progressive neurodegenerative disease. Like for other dementias, biomarkers may help characterize distinct aspects of the underlying pathology, predict decline, and monitor disease progression. We extracted a composite array of 24 speech‐based biomarkers using machine learning and signal processing techniques. We then determined which of the given biomarkers was predictive of MCI and AD progression from baseline by conducting a correlation analysis. Method Spoken responses to a Cookie Theft Picture description task were recorded from 2 subjects with AD, 4 subjects with Mild Cognitive Impairment (MCI) due to AD (biomarker confirmed AD), and 4 subjects with MCI. We automatically extracted acoustic, linguistic, and cognitive biomarkers using speech recognition, acoustic, and language modeling techniques. We later computed Kendall’s tau‐b ( τ b ) correlation between the biomarker values and Montreal Cognitive Assessment (MoCA) scores. Recordings and MoCA scores were collected at baseline, at 6 months after baseline, and at 12 months for a subgroup of the subjects. To perform the correlation analysis, we used scipy.stats.kendalltau library in Python. Result With respect to the acoustic biomarkers, a strong negative correlation ( τ b = ‐0.50) was observed between the MoCA scores and features such as pause time, pause percentage, and pause speech ratio, while a moderate positive correlation ( τ b = 0.2) was found for rhythm standard deviation. With respect to the linguistic biomarkers a strong positive correlation( τ b = 0.50) for the features of corrected type‐token ratio, average sentence length in words, and noun count was reported, while a moderate positive correlation ( τ b = 0.27) was found for the features of word type count, word token count, and moving average type‐token ratio. The other acoustic, linguistic, and cognitive biomarkers showed a very weak ( τ b = 0.00 ‐ 0.10) or weak correlation ( τ b = 0.10 ‐ 0.20). Conclusion Altogether, these results suggest that speech‐based interpretable biomarkers may help clinicians to diagnose AD at earlier stages and monitor disease progression. Our preliminary data suggest that AD patients encounter more problems delivering longer and linguistically elaborated narratives as the disease progresses.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".