Objective speech measures capture depressive symptoms and associated cognitive difficulties
Bibliographic record
Abstract
Psychiatry lacks objective biomarkers for assessing depression, relying instead on subjective measures, such as the Hamilton Depression Rating Scale (HAMD-17). This study examined whether speech features could serve as objective markers of depressive symptoms and its associated cognitive difficulties. Sixty-six individuals with major depressive disorder (MDD) and 54 non-depressed control participants completed a speech assessment, responding to the prompt: "Please tell me how you are feeling today." Linguistic (valence, emotional intensity, agency) and acoustic (pitch, pitch variance, speech rate, time spent pausing) features were derived from natural language processing. These speech features were analyzed individually and collectively as a composite score representing overall speech disturbance. A subset of participants (40 with MDD, 38 controls) also completed a validated executive function task. ANCOVA models compared speech features between groups. Linear regression models examined associations between speech features, depression severity (HAMD-17), and performance on an executive function task. Compared to controls, individuals with MDD used language that was more negatively valenced, emotionally intense, and less agentic. They also demonstrated lower pitch, slower speech rate, and more time spent pausing. The composite speech score also differed between groups. Speech features and executive function were not associated with depression severity, as measured by the HAMD-17. However, several speech features were associated with executive function. Taken together, these findings suggest that speech features may provide a scalable, objective method for detecting depressive symptoms and associated executive difficulties.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".