Tracking and Predicting Depressive Symptoms of Adolescents Using Smartphone-Based Self-Reports, Parental Evaluations, and Passive Phone Sensor Data: Development and Usability Study
Bibliographic record
Abstract
BACKGROUND: Depression carries significant financial, medical, and emotional burden on modern society. Various proof-of-concept studies have highlighted how apps can link dynamic mental health status changes to fluctuations in smartphone usage in adult patients with major depressive disorder (MDD). However, the use of such apps to monitor adolescents remains a challenge. OBJECTIVE: This study aimed to investigate whether smartphone apps are useful in evaluating and monitoring depression symptoms in a clinically depressed adolescent population compared with the following gold-standard clinical psychometric instruments: Patient Health Questionnaire (PHQ-9), Hamilton Rating Scale for Depression (HAM-D), and Hamilton Anxiety Rating Scale (HAM-A). METHODS: We recruited 13 families with adolescent patients diagnosed with MDD with or without comorbid anxiety disorder. Over an 8-week period, daily self-reported moods and smartphone sensor data were collected by using the Smartphone- and OnLine usage-based eValuation for Depression (SOLVD) app. The evaluations from teens' parents were also collected. Baseline depression and anxiety symptoms were measured biweekly using PHQ-9, HAM-D, and HAM-A. RESULTS: We observed a significant correlation between the self-evaluated mood averaged over a 2-week period and the biweekly psychometric scores from PHQ-9, HAM-D, and HAM-A (0.45≤|r|≤0.63; P=.009, P=.01, and P=.003, respectively). The daily steps taken, SMS frequency, and average call duration were also highly correlated with clinical scores (0.44≤|r|≤0.72; all P<.05). By combining self-evaluations and smartphone sensor data of the teens, we could predict the PHQ-9 score with an accuracy of 88% (23.77/27). When adding the evaluations from the teens' parents, the prediction accuracy was further increased to 90% (24.35/27). CONCLUSIONS: Smartphone apps such as SOLVD represent a useful way to monitor depressive symptoms in clinically depressed adolescents, and these apps correlate well with current gold-standard psychometric instruments. This is a first study of its kind that was conducted on the adolescent population, and it included inputs from both teens and their parents as observers. The results are preliminary because of the small sample size, and we plan to expand the study to a larger population.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".