Gender Differences in Trajectories of Depressive Symptoms Among Talkspace Clients: Naturalistic Observational Study
Bibliographic record
Abstract
Background: Gender minority populations experience an increased risk of depression and report significant barriers to accessing mental health services. While digital mental health (DMH) technologies may address barriers, it remains unclear how gender minority clients engage with DMH services and if DMH improves their clinical outcomes. Objective: This naturalistic study explored gender differences in 15-week clinical outcomes of clients receiving technology-mediated psychotherapy from a large DMH provider. Methods: This study used observational data of clients who signed up for Talkspace (Talkspace, Inc) between February 2017 and July 2021. The analytic sample included Talkspace clients (N=20,156) with a baseline 8-item Patient Health Questionnaire (PHQ-8) score ≥10. Participants completed at least 2 PHQ-8 assessments over 15 weeks of treatment. Multilevel linear models tested gender differences in depressive symptom trajectories over the course of treatment (model 1) while also controlling for baseline PHQ-8 scores (model 2) and treatment engagement indicators (model 3). Sensitivity analyses reestimated model 2 among clients who submitted a PHQ-8 survey during the week 15 assessment period and among those who discontinued treatment beforehand. Reasons for service cancellation were also described for the latter group. Gender differences in secondary clinical outcomes were examined via chi-square and Fisher exact tests. Results: In all models, there were significant week-by-gender interactions. When controlling for baseline PHQ-8 scores, rates of symptom change were significantly slower for gender-diverse participants (b=0.60; P<.001), nonbinary participants (b=0.81; P<.001), and transgender women (b=0.87; P=.007), but not for women (P=.98) or transgender men (P=.38) compared to men. By week 15, adjusted PHQ-8 scores declined 8.7 points for both men and women, versus 4.4-7.4 points for gender minority clients. Sensitivity analyses indicated attenuated symptom improvement among week-15 completers, with transgender women showing the slowest changes (b=0.76; P=.02). Among earlier dropouts, weekly symptom reductions were steep overall (eg, week 3: b=-4.06, P<.001; week 6: b=-2.31, P<.001) while certain gender minority subgroups worsened (eg, adjusted scores for transgender women increased from 15.41 at baseline to 16.08 at final week 3 PHQ-8 survey submissions). Cancellation data (3450/20,156, 17.12%) confirmed discontinuation reasons related to both symptom improvement (928/3691 reasons, 25.14%) and potential barriers to treatment engagement (eg, cost: 1431/3691, 38.77%; poor service fit or poor perceived effectiveness: 677/3691, 18.34%). Gender differences were observed in rates of treatment response (weeks 3-12; all P≤.02), symptom remission (weeks 3, 6, 9, and 15; all P≤.047), and clinically significant symptom reduction (all time points, all P≤.03). Symptom deterioration did not differ by gender (all P>.05). Conclusions: While clinical outcomes generally improved over time among clients engaged in technology-mediated psychotherapy, some gender minority populations experienced slower improvements. Future research may explore strategies to adapt DMH interventions to better meet the needs of diverse gender identities.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".