Bias in Acoustic‐ and Linguistic‐based Classifications of Alzheimer’s Disease
Bibliographic record
Abstract
Abstract Background It is crucial to identify patients with Alzheimer’s disease (AD). To do so, various attempts have been made to develop AI and non‐AI assessments, among which some focus on developing acoustic‐ and linguistic‐based classifiers. They are based on the detection of impairments in their language and speech, which can manifest years before other cognitive impairments associated with AD appear. The fact is that most of the current vocal and language classifiers of AD have been trained using the Pitt corpus, which is an imbalanced class and gender dataset. These two characteristics could be enough to bias such classifiers making them untrustworthy and unsuitable for integration into AD care settings. This paper presents a novel method for collecting vocal data to reduce the impact of potential sources of bias on acoustic and linguistic classifiers for AD. Method We will use a data collection approach that collects voices from participants (i.e., patients with AD and healthy controls) with diversity in race, gender, educational, socioeconomic, and cultural backgrounds. Participants will be asked to perform diverse language tasks, including word association, verbal fluency, and grammaticality judgment tasks. Following such an approach can ensure we collect oral data that wouldn’t be affected by selection, recruitment, sociolinguistics, and gender biases. Results The main result of this study is the creation of a benchmark vocal dataset from diverse gender and racial groups and ensuring that the data is representative of the population with AD. We expect such data to be used for evaluating and validating and help AI developers to successfully develop fair acoustic and linguistic classifiers of AD. It, in its turn, can motivate healthcare professionals to employ these systems as assistants to identify patients with AD from their voices quickly. Conclusion The Pitt corpus is a biased data set. Thus, acoustic and linguistic classifiers trained upon it can not be considered trustworthy AI systems. The AI developers that aim to deploy vocal systems into Alzheimer’s disease care settings would need unbiased vocal data. This study proposed a method to collect such verbal data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.054 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".