Predicting Depressive Symptom Severity Through Individuals’ Nearby Bluetooth Device Count Data Collected by Mobile Phones: Preliminary Longitudinal Study
Bibliographic record
Abstract
Background Research in mental health has found associations between depression and individuals’ behaviors and statuses, such as social connections and interactions, working status, mobility, and social isolation and loneliness. These behaviors and statuses can be approximated by the nearby Bluetooth device count (NBDC) detected by Bluetooth sensors in mobile phones. Objective This study aimed to explore the value of the NBDC data in predicting depressive symptom severity as measured via the 8-item Patient Health Questionnaire (PHQ-8). Methods The data used in this paper included 2886 biweekly PHQ-8 records collected from 316 participants recruited from three study sites in the Netherlands, Spain, and the United Kingdom as part of the EU Remote Assessment of Disease and Relapse-Central Nervous System (RADAR-CNS) study. From the NBDC data 2 weeks prior to each PHQ-8 score, we extracted 49 Bluetooth features, including statistical features and nonlinear features for measuring the periodicity and regularity of individuals’ life rhythms. Linear mixed-effect models were used to explore associations between Bluetooth features and the PHQ-8 score. We then applied hierarchical Bayesian linear regression models to predict the PHQ-8 score from the extracted Bluetooth features. Results A number of significant associations were found between Bluetooth features and depressive symptom severity. Generally speaking, along with depressive symptom worsening, one or more of the following changes were found in the preceding 2 weeks of the NBDC data: (1) the amount decreased, (2) the variance decreased, (3) the periodicity (especially the circadian rhythm) decreased, and (4) the NBDC sequence became more irregular. Compared with commonly used machine learning models, the proposed hierarchical Bayesian linear regression model achieved the best prediction metrics (R2=0.526) and a root mean squared error (RMSE) of 3.891. Bluetooth features can explain an extra 18.8% of the variance in the PHQ-8 score relative to the baseline model without Bluetooth features (R2=0.338, RMSE=4.547). Conclusions Our statistical results indicate that the NBDC data have the potential to reflect changes in individuals’ behaviors and statuses concurrent with the changes in the depressive state. The prediction results demonstrate that the NBDC data have a significant value in predicting depressive symptom severity. These findings may have utility for the mental health monitoring practice in real-world settings.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".