Passively Captured Interpersonal Social Interactions and Motion From Smartphones for Predicting Decompensation in Heart Failure: Observational Cohort Study
Bibliographic record
Abstract
BACKGROUND: Heart failure (HF) is a major cause of frequent hospitalization and death. Early detection of HF symptoms using smartphone-based monitoring may reduce adverse events in a low-cost, scalable way. OBJECTIVE: We examined the relationship of HF decompensation events with smartphone-based features derived from passively and actively acquired data. METHODS: This was a prospective cohort study in which we monitored HF participants' social and movement activities using a smartphone app and followed them for clinical events via phone and chart review and classified the encounters as compensated or decompensated by reviewing the provider notes in detail. We extracted motion, location, and social interaction passive features and self-reported quality of life weekly (active) with the short Kansas City Cardiomyopathy Questionnaire (KCCQ-12) survey. We developed and validated an algorithm for classifying decompensated versus compensated clinical encounters (hospitalizations or clinic visits). We evaluated models based on single modality as well as early and late fusion approaches combining patient-reported outcomes and passive smartphone data. We used Shapley additive explanation values to quantify the contribution and impact of each feature to the model. RESULTS: We evaluated 28 participants with a mean age of 67 years (SD 8), among whom 11% (3/28) were female and 46% (13/28) were Black. We identified 62 compensated and 48 decompensated clinical events from 24 and 22 participants, respectively. The highest area under the precision-recall curve (AUCPr) for classifying decompensation was with a late fusion approach combining KCCQ-12, motion, and social contact features using leave-one-subject-out cross-validation for a 2-day prediction window. It had an AUCPr of 0.80, with an area under the receiver operator curve (AUC) of 0.83, a positive predictive value (PPV) of 0.73, a sensitivity of 0.77, and a specificity of 0.88 for a 2-day prediction window. Similarly, the 4-day window model had an AUC of 0.82, an AUCPr of 0.69, a PPV of 0.62, a sensitivity of 0.68, and a specificity of 0.87. Passive social data provided some of the most informative features, with fewer calls of longer duration associating with a higher probability of future HF decompensation. CONCLUSIONS: Smartphone-based data that includes both passive monitoring and actively collected surveys may provide important behavioral and functional health information on HF status in advance of clinical visits. This proof-of-concept study, although small, offers important insight into the social and behavioral determinants of health and the feasibility of using smartphone-based monitoring in this population. Our strong results are comparable to those of more active and expensive monitoring approaches, and underscore the need for larger studies to understand the clinical significance of this monitoring method.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".