Passive Sensor Data for Characterizing States of Increased Risk for Eating Disorder Behaviors in the Digital Phenotyping Arm of the Binge Eating Genetics Initiative: Protocol for an Observational Study
Bibliographic record
Abstract
BACKGROUND: Data that can be easily, efficiently, and safely collected via cell phones and other digital devices have great potential for clinical application. Here, we focus on how these data could be used to refine and augment intervention strategies for binge eating disorder (BED) and bulimia nervosa (BN), conditions that lack highly efficacious, enduring, and accessible treatments. These data are easy to collect digitally but are highly complex and present unique methodological challenges that invite innovative solutions. OBJECTIVE: We describe the digital phenotyping component of the Binge Eating Genetics Initiative, which uses personal digital device data to capture dynamic patterns of risk for binge and purge episodes. Characteristic data signatures will ultimately be used to develop personalized models of eating disorder pathologies and just-in-time interventions to reduce risk for related behaviors. Here, we focus on the methods used to prepare the data for analysis and discuss how these approaches can be generalized beyond the current application. METHODS: The University of North Carolina Biomedical Institutional Review Board approved all study procedures. Participants who met diagnostic criteria for BED or BN provided real time assessments of eating behaviors and feelings through the Recovery Record app delivered on iPhones and the Apple Watches. Continuous passive measures of physiological activation (heart rate) and physical activity (step count) were collected from Apple Watches over 30 days. Data were cleaned to account for user and device recording errors, including duplicate entries and unreliable heart rate and step values. Across participants, the proportion of data points removed during cleaning ranged from <0.1% to 2.4%, depending on the data source. To prepare the data for multivariate time series analysis, we used a novel data handling approach to address variable measurement frequency across data sources and devices. This involved mapping heart rate, step count, feeling ratings, and eating disorder behaviors onto simultaneous minute-level time series that will enable the characterization of individual- and group-level regulatory dynamics preceding and following binge and purge episodes. RESULTS: Data collection and cleaning are complete. Between August 2017 and May 2021, 1019 participants provided an average of 25 days of data yielding 3,419,937 heart rate values, 1,635,993 step counts, 8274 binge or purge events, and 85,200 feeling observations. Analysis will begin in spring 2022. CONCLUSIONS: We provide a detailed description of the methods used to collect, clean, and prepare personal digital device data from one component of a large, longitudinal eating disorder study. The results will identify digital signatures of increased risk for binge and purge events, which may ultimately be used to create digital interventions for BED and BN. Our goal is to contribute to increased transparency in the handling and analysis of personal digital device data. TRIAL REGISTRATION: ClinicalTrials.gov NCT04162574; https://clinicaltrials.gov/ct2/show/NCT04162574. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): DERR1-10.2196/38294.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.033 | 0.032 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.005 | 0.003 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.003 | 0.005 |
| Insufficient payload (model declined to judge) | 0.037 | 0.012 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".