Reliability and generalizability of neural speech tracking in younger and older adults
Bibliographic record
Abstract
Abstract Neural tracking of continuous, spoken speech is increasingly used to examine how the brain encodes speech and is considered a potential clinical biomarker, for example, for age-related hearing loss. A biomarker must be reliable (intra-class correlation [ICC] >0.7), but the reliability of neural-speech tracking is unclear. In the current study, younger and older adults (different genders) listened to stories in two separate sessions while electroencephalography (EEG) was recorded in order to investigate the reliability and generalizability of neural speech tracking. Neural speech tracking was larger for older compared to younger adults for stories under clear and background noise conditions, consistent with a loss of inhibition in the aged auditory system. For both age groups, reliability for neural speech tracking was lower than the reliability of neural responses to noise bursts (ICC >0.8), which we used as a benchmark for maximum reliability. The reliability of neural speech tracking was moderate (ICC ∼0.5-0.75) but tended to be lower for younger adults when speech was presented in noise. Neural speech tracking also generalized moderately across different stories (ICC ∼0.5-0.6), which appeared greatest for audiobook-like stories spoken by the same person. This indicates that a variety of stories could possibly be used for clinical assessments. Overall, the current data provide results critical for the development of a biomarker of speech processing, but also suggest that further work is needed to increase the reliability of the neural-tracking response to meet clinical standards. Significance statement Neural speech tracking approaches are increasingly used in research and considered a biomarker for impaired speech processing. A biomarker needs to be reliable, but the reliability of neural speech tracking is unclear. The current study shows in younger and older adults that the neural-tracking response is moderately reliable (ICC ∼0.5-0.75), although more variable in younger adults, and that the tracking response also moderately generalize across different stories (ICC ∼0.5-0.6), especially for audiobook-like stories spoken by the same person. The current data provide results critical for the development of a biomarker of speech processing, but also suggest that further work is needed to increase the reliability of the neural-tracking response to meet clinical standards.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.034 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".