Exploring Digital Biomarkers of Illness Activity in Mood Episodes: Hypotheses Generating and Model Development Study
Bibliographic record
Abstract
BACKGROUND: Depressive and manic episodes within bipolar disorder (BD) and major depressive disorder (MDD) involve altered mood, sleep, and activity, alongside physiological alterations wearables can capture. OBJECTIVE: Firstly, we explored whether physiological wearable data could predict (aim 1) the severity of an acute affective episode at the intra-individual level and (aim 2) the polarity of an acute affective episode and euthymia among different individuals. Secondarily, we explored which physiological data were related to prior predictions, generalization across patients, and associations between affective symptoms and physiological data. METHODS: We conducted a prospective exploratory observational study including patients with BD and MDD on acute affective episodes (manic, depressed, and mixed) whose physiological data were recorded using a research-grade wearable (Empatica E4) across 3 consecutive time points (acute, response, and remission of episode). Euthymic patients and healthy controls were recorded during a single session (approximately 48 h). Manic and depressive symptoms were assessed using standardized psychometric scales. Physiological wearable data included the following channels: acceleration (ACC), skin temperature, blood volume pulse, heart rate (HR), and electrodermal activity (EDA). Invalid physiological data were removed using a rule-based filter, and channels were time aligned at 1-second time units and segmented at window lengths of 32 seconds, as best-performing parameters. We developed deep learning predictive models, assessed the channels' individual contribution using permutation feature importance analysis, and computed physiological data to psychometric scales' items normalized mutual information (NMI). We present a novel, fully automated method for the preprocessing and analysis of physiological data from a research-grade wearable device, including a viable supervised learning pipeline for time-series analyses. RESULTS: Overall, 35 sessions (1512 hours) from 12 patients (manic, depressed, mixed, and euthymic) and 7 healthy controls (mean age 39.7, SD 12.6 years; 6/19, 32% female) were analyzed. The severity of mood episodes was predicted with moderate (62%-85%) accuracies (aim 1), and their polarity with moderate (70%) accuracy (aim 2). The most relevant features for the former tasks were ACC, EDA, and HR. There was a fair agreement in feature importance across classification tasks (Kendall W=0.383). Generalization of the former models on unseen patients was of overall low accuracy, except for the intra-individual models. ACC was associated with "increased motor activity" (NMI>0.55), "insomnia" (NMI=0.6), and "motor inhibition" (NMI=0.75). EDA was associated with "aggressive behavior" (NMI=1.0) and "psychic anxiety" (NMI=0.52). CONCLUSIONS: Physiological data from wearables show potential to identify mood episodes and specific symptoms of mania and depression quantitatively, both in BD and MDD. Motor activity and stress-related physiological data (EDA and HR) stand out as potential digital biomarkers for predicting mania and depression, respectively. These findings represent a promising pathway toward personalized psychiatry, in which physiological wearable data could allow the early identification and intervention of mood episodes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".