MétaCan
Menu
Back to cohort
Record W3102950836 · doi:10.1016/j.eclinm.2020.100644

Towards early prediction of Alzheimer's disease through language samples

2020· article· en· W3102950836 on OpenAlexaffabout
Jed A. Meltzer

Bibliographic record

VenueEClinicalMedicine · 2020
Typearticle
Languageen
FieldNeuroscience
TopicNeurobiology of Language and Bilingualism
Canadian institutionsBaycrest HospitalUniversity of Toronto
Fundersnot available
KeywordsPrimary progressive aphasiaMedicineScopusDiseaseNeuropsychologyAphasiaDementiaNeurodegenerationAlzheimer's diseaseBiomarkerCognitive psychologyMEDLINEPsychologyPathologyPsychiatryCognitionFrontotemporal dementia

Abstract

fetched live from OpenAlex

Although accurate diagnosis of Alzheimer's Disease (AD) remains a priority for research, even more research interest currently focuses on the prediction of the disease years or decades before its onset. Because the neurodegeneration caused by the disease is likely irreversible, a better treatment strategy would be to identify those undergoing the early changes linked to eventual disease onset and to administer a mitigating treatment (yet to be developed) at that time. One biomarker of intense interest is naturalistic language samples, as they are easy to acquire, completely noninvasive, and, compared to most neuropsychological assessments, easily repeated on a regular basis without practice effects. However, the analysis is complicated, laborious, and potentially subjective. In recent years, advances in machine learning and natural language processing have been applied to language samples for the detection of dementia, and researchers have achieved considerable success in distinguishing the speech of individuals with and without dementia [1Orimaye S.O. Wong J.S. Golden K.J. Wong C.P. Soyiri I.N Predicting probable Alzheimer's disease using linguistic deficits and biomarkers.BMC Bioinformatics. 2017; 18: 34Crossref PubMed Scopus (61) Google Scholar, 2Fraser K.C. Meltzer J.A. Rudzicz F Linguistic features identify alzheimer's disease in narrative speech.J Alzheimers Dis. 2015; 49: 407-422Crossref Scopus (305) Google Scholar, 3Fraser K.C. Meltzer J.A. Graham N.L. Leonard C. Hirst G. Black S.E. et al.Automated classification of primary progressive aphasia subtypes from narrative speech transcripts.Cortex. 2014; 55: 43-60Summary Full Text Full Text PDF PubMed Scopus (118) Google Scholar]. Despite these advances, the predictive power of language samples is largely unproven, given that very few studies have been able to examine participants years before an eventual diagnosis of AD, to compare the language output of those who do and do not go on to develop the disease [[4]Snowdon D.A. Kemper S.J. Mortimer J.A. Greiner L.H. Wekstein D.R. Markesbery W.R Linguistic ability in early life and cognitive function and Alzheimer's disease in late life. Findings from the Nun Study.JAMA. 1996; 275: 528-532Crossref PubMed Google Scholar]. A prospective study of this topic would require a very large sample, take many years to complete, and would have a relatively low yield of positive findings for the effort. Fortunately, as the potential value of speech and other cognitive measures as a biomarker has come to increased attention, large-scale prospective studies of health in general have begun to include them in their assessment batteries. As published in EClinicalMedicine, Elif Eyigoz and colleagues present an analysis [[5]Eyigoz E. Mathur S. Santamaria M. Cecchi G. Naylor M Linguistic markers predict onset of Alzheimer's disease.Eclinical Medicine. 2020; https://doi.org/10.1016/j.eclinm.2020.100583Summary Full Text Full Text PDF PubMed Scopus (28) Google Scholar] of written language samples collected in the Framingham Heart Study (FHS), one of the world's largest and best-known prospective health studies. Founded in 1948, the FHS began to incorporate a neuropsychological test battery in 1999, including a brief written picture description. Critically, in the years since this was introduced, some participants went on to develop dementia while many more did not, allowing a retrospective comparison based on data fortuitously acquired years earlier. Elif Eyigoz et al. applied state of the art analysis procedures to samples from 270 participants finding that dementia onset before age 85 could be identified with 70–75% accuracy. This is well above chance performance, and slightly below the best performance seen in studies comparing current dementia patients with controls [[6]Petti U. Baker S. Korhonen A A systematic literature review of automatic Alzheimer's disease detection from speech and language.J Am Med Inform Assoc. 2020; Crossref PubMed Scopus (18) Google Scholar]. The authors then demonstrate that their model picks up on many of the same linguistic trends seen in previous comparisons of AD vs. controls. This finding is exciting, being among the first to show predictive value of language samples well before the onset of dementia, while its limitations point the way for future work. The accuracy rates of 70–75% are encouraging but not yet satisfactory for a realistic clinical tool, but this is likely to be an inevitable consequence of the limited language samples available. The samples are descriptions of a single picture (“Cookie Theft”), and limited both in length to a few dozen words, and in content to what is shown in the picture. Longer samples covering more extensive topics would surely provide a more sensitive view of linguistic changes. More significantly, the samples are only written. Although written samples have a history of predictive value, speech is a more natural form of communication giving access to several important quantitative variables, especially those related to the ease of word finding, including overall speech rate and pausing [[7]Guo Z. Ling Z. Li Y Detecting alzheimer's disease from continuous speech using language models.J Alzheimers Dis. 2019; 70: 1163-1174Crossref PubMed Scopus (11) Google Scholar]. It remains to be seen how high the predictive accuracy of language samples can be pushed given more extensive data sources, and it is important that such data be collected. The most advanced machine learning algorithms cannot overcome the limitations of sparse input. Fortunately, collection of speech data is simple to implement and simple to include within a larger natural history study like the FHS. A number of longitudinal health studies now include detailed speech measures and can be expected to yield new insights into the earliest stages of the evolution of dementia [[8]Mueller K.D. Koscik R.L. Hermann B.P. Johnson S.C. Turkstra L.S Declines in connected language are associated with very early mild cognitive impairment: results from the wisconsin registry for alzheimer's prevention.Front Aging Neurosci. 2018; 9PubMed Google Scholar,[9]Farhan S.M. Bartha R. Black S.E. Corbett D. Finger E. Freedman M. et al.The ontario neurodegenerative disease research initiative (ONDRI).Can J Neurol Sci. 2017; 44: 196-202Crossref PubMed Scopus (34) Google Scholar]. Above all, this study illustrates the value of open science – a simple measure that was not the original focus of the FHS provided valuable new knowledge when shared with the larger scientific community. Given the complexity and expense of longitudinal studies of dementia, this level of data sharing needs to become the norm. Researchers should take pains to harmonize data collection and archiving procedures, while addressing practical concerns such including consent and privacy (especially with speech, as the human voice is inherently identifiable). Additionally, attention should be paid as to which kinds of speech elicitation tasks provide the most informative samples, as there are many different options, including picture description, story retell (e.g. Cinderella), autobiographical interviews, and dyadic conversation [[10]Boschi V. Catricala E. Consonni M. Chesi C. Moro A. Cappa S.F Connected speech in neurodegenerative language disorders: a review.Front Psychol. 2017; 8: 269Crossref PubMed Scopus (107) Google Scholar]. Dr. Meltzer is a shareholder and advisor of Winterlight Labs. Linguistic markers predict onset of Alzheimer's diseaseThe results suggest that language performance in naturalistic probes expose subtle early signs of progression to AD in advance of clinical diagnosis of impairment. Full-Text PDF Open Access

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.004
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: Bench or experimental
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.157
Threshold uncertainty score0.489

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.004
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.144
GPT teacher head0.364
Teacher spread0.220 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations6
Published2020
Admission routes2
Has abstractyes

Explore more

Same venueEClinicalMedicineSame topicNeurobiology of Language and BilingualismFrench-language works237,207