Machine learning for predicting cognitive deficits using auditory and demographic factors
Bibliographic record
Abstract
IMPORTANCE: Predicting neurocognitive deficits using complex auditory assessments could change how cognitive dysfunction is identified, and monitored over time. Detecting cognitive impairment in people living with HIV (PLWH) is important for early intervention, especially in low- to middle-income countries where most cases exist. Auditory tests relate to neurocognitive test results, but the incremental predictive capability beyond demographic factors is unknown. OBJECTIVE: Use machine learning to predict neurocognitive deficits, using auditory tests and demographic factors. SETTING: The Infectious Disease Center in Dar es Salaam, Tanzania. PARTICIPANTS: Participants were 939 Tanzanian individuals from Dar es Salaam living with and without HIV who were part of a longitudinal study. Patients who had only one visit, a positive history of ear drainage, concussion, significant noise or chemical exposure, neurological disease, mental illness, or exposure to ototoxic antibiotics (e.g., gentamycin), or chemotherapy were excluded. This provided 478 participants (349 PLWH, 129 HIV-negative). Participant data were randomized to training and test sets for machine learning. MAIN OUTCOME(S) AND MEASURE(S): The main outcome was whether auditory variables combined with relevant demographic variables could predict neurocognitive dysfunction (defined as a score of <26 on the Kiswahili Montreal Cognitive Assessment) better than demographic factors alone. The performance of predictive machine learning algorithms was primarily evaluated using the area under the receiver operational characteristic curve. Secondary metrics for evaluation included F1 scores, accuracies, and the Youden's indices for the algorithms. RESULTS: The percentage of individuals with cognitive deficits was 36.2% (139 PLWH and 34 HIV-negative). The Gaussian and kernel naïve Bayes classifiers were the most predictive algorithms for neurocognitive impairment. Algorithms trained with auditory variables had average area under the curve values of 0.91 and 0.87, F1 scores (metric for precision and recall) of 0.81 and 0.76, and average accuracies of 86.3% and 81.9% respectively. Algorithms trained without auditory variables as features were statistically worse (p < .001) in both the primary measure of area under the curve (0.82/0.78) and the secondary measure of accuracy (72.3%/74.5%) for the Gaussian and kernel algorithms respectively. CONCLUSIONS AND RELEVANCE: Auditory variables improved the prediction of cognitive function. Since auditory tests are easy-to-administer and often naturalistic tasks, they may offer objective measures or predictors of neurocognitive performance suitable for many global settings. Further research and development into using machine learning algorithms for predicting cognitive outcomes should be pursued.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.017 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".