Testing the Pragmatic Effectiveness of a Consumer-Based Mindfulness Mobile App in the Workplace: Randomized Controlled Trial
Bibliographic record
Abstract
BACKGROUND: Mental health and sleep problems are prevalent in the workforce, corresponding to costly impairment in productivity and increased health care use. Digital mindfulness interventions are efficacious in improving sleep and mental health in the workplace; however, evidence supporting their pragmatic utility, potential for improving productivity, and ability to reduce employer costs is limited. OBJECTIVE: This pragmatic, cluster randomized controlled trial aimed to evaluate the experimental effects of implementing a commercially available mindfulness app-Calm-in employees of a large, multisite employer in the United States. Outcomes included mental health (depression, anxiety, and stress), sleep (insomnia and daytime sleepiness), resilience, productivity impairment (absenteeism, presenteeism, overall work impairment, and non-work activity impairment), and health care use (medical visit frequency). METHODS: Employees were randomized at the work site to receive either the Calm app intervention or waitlist control. Participants in the Calm intervention group were instructed to use the Calm app for 10 minutes per day for 8 weeks; individuals with elevated baseline insomnia symptoms could opt-in to 6 weeks of sleep coaching. All outcomes were assessed every 2 weeks, with the exception of medical visits (weeks 4 and 8 only). Effects of the Calm intervention on outcomes were evaluated via mixed effects modeling, controlling for relevant baseline characteristics, with fixed effects of the intervention on outcomes assessed at weeks 2, 4, 6, and 8. Models were analyzed via complete-case and intent-to-treat analyses. RESULTS: A total of 1029 employees enrolled (n=585 in the Calm intervention group, including 101 who opted-in to sleep coaching, and n=444 in waitlist control). Of them, 192 (n=88 for the Calm intervention group and n=104 for waitlist) completed all 5 assessments. In the complete-case analysis at week 8, employees at sites randomized to the Calm intervention group experienced significant improvements in depression (P=.02), anxiety (P=.01), stress (P<.001), insomnia (P<.001), sleepiness (P<.001), resilience (P=.02), presenteeism (P=.01), overall work impairment (P=.004), and nonwork impairment (P<.001), and reduced medical care visit frequency (P<.001) and productivity impairment costs (P=.01), relative to the waitlist control. In the intent-to-treat analysis at week 8, significant benefits of the intervention were observed for depression (P=.046), anxiety (P=.01), insomnia (P<.001), sleepiness (P<.001), nonwork impairment (P=.04), and medical visit frequency (P<.001). CONCLUSIONS: The results suggest that the Calm app is an effective workplace intervention for improving mental health, sleep, resilience, and productivity and for reducing medical visits and costs owing to work impairment. Future studies should identify optimal implementation strategies that maximize employee uptake and large-scale implementation success across diverse, geographically dispersed employers. TRIAL REGISTRATION: ClinicalTrials.gov NCT05120310; https://clinicaltrials.gov/ct2/show/NCT05120310.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".