Computer-based characterization of language alterations throughout the Alzheimer's disease continuum
Notice bibliographique
Résumé
According to the American and Canadian Alzheimer’s Associations, research into methods for the early detection of Alzheimer’s disease is imperative. Many studies have emphasized the numerous advantages for patients, family members and governments of detecting the disease at the pre-clinical stage of its continuum. However, at this stage, changes are very subtle, making their detection a challenging task. \n \nAlterations in language functions have been found years before the dementia stage of the disease continuum. For this reason, many researchers have focused their efforts on investigating methods for identifying cues of the presence of the disease hidden in language. \n \nOne type of cognitive test commonly used in this type of research consists of standardized picture description tasks. These tasks elicit the speech of patients through a visual stimulus, and are usually part of cognitive assessment batteries used in clinical practice. The tasks have the advantage of presenting patients with a single constrained thematic, which limits the vocabulary and facilitates comparisons across patients and languages. However, they also limit the variety of syntactic structures, hindering some linguistic analyses, and being a part of usual clinical examinations, may increase nervousness in some patients. \n \nThe study of spontaneous conversations is an alternative to using picture description tasks for language analyses. Spontaneous conversations have the advantage of allowing the use of unconstrained idiosyncratic syntactic structures and vocabulary. They are also less stressful to patients and could be conducted with a nurse, a caregiver or a person familiar to the patient. Nevertheless, many factors, such as socio-demographic and cultural differences, may define the linguistic characteristics of individuals. Consequently, a characterization of the changes in language functions that occur during the continuum of the disease could be helpful in the monitoring of patient-specific changes. \n \nThis doctoral thesis presents a computer-based methodology for evaluating patients’ performance during standardized picture description tasks, and for assessing language functions in the context of these tasks and in spontaneous conversations. We believe that both evaluations can complement each other and provide an inexpensive and noninvasive method for monitoring language functions. In practice, picture description tasks could be realized routinely at the doctor’s office, while spontaneous conversations could be held at more regular intervals and at more convenient locations for the patient. \n \nFor our work, we compared the computed performance and language functions of patients during standardized picture description tasks against a population with similar socio-demographic characteristics. For this, our proposed method evaluated the informativeness and pertinence of the descriptions of patients, as well as their lexical richness. Using our metrics, we trained machine learning algorithms to estimate their adeptness at differentiating Alzheimer’s patients from healthy controls. We obtained an area under the curve of 0.83 in this task. We also achieved an area under the curve of 0.79 for classifying healthy controls and patients with mild cognitive impairment, which is often a pre-clinal precursor of Alzheimer’s disease. \n \nIn addition, we proposed an automated method for evaluating lexical richness, vocabulary distribution, speech fluidity and the use of specific syntactic structures among older French speakers during spontaneous conversations. We characterized the changes that four speakers underwent as they transitioned from a healthy state to some form of cognitive disease, including Alzheimer’s disease. We observed marked differences in our proposed metrics between those individuals that would develop a cognitive disease and healthy matched controls, even when analyzing transcriptions of conversations from up to ten years before the time of diagnosis. \n \nAs a concomitant contribution of this doctoral work, we designed the protocol and created the Spanish cohort of the Carolinas’ Conversations Collection. This cohort includes longitudinal video-recordings and transcriptions of spontaneous conversations of older Spanish speakers in Mexico and Ecuador. These recollections are the result of the combined efforts of six institutions from four different countries, and will be available for research purposes upon request. This undertaking is aimed at lessening the scarcity of data of this type, and at encouraging research on language and communication in the older population.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,003 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,004 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,001 | 0,001 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,003 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».