Alignment Between Heart Rate Variability From Fitness Trackers and Perceived Stress: Perspectives From a Large-Scale In Situ Longitudinal Study of Information Workers
Bibliographic record
Abstract
Background Stress can have adverse effects on health and well-being. Informed by laboratory findings that heart rate variability (HRV) decreases in response to an induced stress response, recent efforts to monitor perceived stress in the wild have focused on HRV measured using wearable devices. However, it is not clear that the well-established association between perceived stress and HRV replicates in naturalistic settings without explicit stress inductions and research-grade sensors. Objective This study aims to quantify the strength of the associations between HRV and perceived daily stress using wearable devices in real-world settings. Methods In the main study, 657 participants wore a fitness tracker and completed 14,695 ecological momentary assessments (EMAs) assessing perceived stress, anxiety, positive affect, and negative affect across 8 weeks. In the follow-up study, approximately a year later, 49.8% (327/657) of the same participants wore the same fitness tracker and completed 1373 EMAs assessing perceived stress at the most stressful time of the day over a 1-week period. We used mixed-effects generalized linear models to predict EMA responses from HRV features calculated over varying time windows from 5 minutes to 24 hours. Results Across all time windows, the models explained an average of 1% (SD 0.5%; marginal R2) of the variance. Models using HRV features computed from an 8 AM to 6 PM time window (namely work hours) outperformed other time windows using HRV features calculated closer to the survey response time but still explained a small amount (2.2%) of the variance. HRV features that were associated with perceived stress were the low frequency to high frequency ratio, very low frequency power, triangular index, and SD of the averages of normal-to-normal intervals. In addition, we found that although HRV was also predictive of other related measures, namely, anxiety, negative affect, and positive affect, it was a significant predictor of stress after controlling for these other constructs. In the follow-up study, calculating HRV when participants reported their most stressful time of the day was less predictive and provided a worse fit (R2=0.022) than the work hours time window (R2=0.032). Conclusions A significant but small relationship between perceived stress and HRV was found. Thus, although HRV is associated with perceived stress in laboratory settings, the strength of that association diminishes in real-life settings. HRV might be more reflective of perceived stress in the presence of specific and isolated stressors and research-grade sensing. Relying on wearable-derived HRV alone might not be sufficient to detect stress in naturalistic settings and should not be considered a proxy for perceived stress but rather a component of a complex phenomenon.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".