Development of the ehive Digital Health App: Protocol for a Centralized Research Platform
Bibliographic record
Abstract
BACKGROUND: The increasing use of smartphones, wearables, and connected devices has enabled the increasing application of digital technologies for research. Remote digital study platforms comprise a patient-interfacing digital application that enables multimodal data collection from a mobile app and connected sources. They offer an opportunity to recruit at scale, acquire data longitudinally at a high frequency, and engage study participants at any time of the day in any place. Few published descriptions of centralized digital research platforms provide a framework for their development. OBJECTIVE: This study aims to serve as a road map for those seeking to develop a centralized digital research platform. We describe the technical and functional aspects of the ehive app, the centralized digital research platform of the Hasso Plattner Institute for Digital Health at Mount Sinai Hospital, New York, New York. We then provide information about ongoing studies hosted on ehive, including usership statistics and data infrastructure. Finally, we discuss our experience with ehive in the broader context of the current landscape of digital health research platforms. METHODS: The ehive app is a multifaceted and patient-facing central digital research platform that permits the collection of e-consent for digital health studies. An overview of its development, its e-consent process, and the tools it uses for participant recruitment and retention are provided. Data integration with the platform and the infrastructure supporting its operations are discussed; furthermore, a description of its participant- and researcher-facing dashboard interfaces and the e-consent architecture is provided. RESULTS: The ehive platform was launched in 2020 and has successfully hosted 8 studies, namely 6 observational studies and 2 clinical trials. Approximately 1484 participants downloaded the app across 36 states in the United States. The use of recruitment methods such as bulk messaging through the EPIC electronic health records and standard email portals enables broad recruitment. Light-touch engagement methods, used in an automated fashion through the platform, maintain high degrees of engagement and retention. The ehive platform demonstrates the successful deployment of a central digital research platform that can be modified across study designs. CONCLUSIONS: Centralized digital research platforms such as ehive provide a novel tool that allows investigators to expand their research beyond their institution, engage in large-scale longitudinal studies, and combine multimodal data streams. The ehive platform serves as a model for groups seeking to develop similar digital health research programs. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): DERR1-10.2196/49204.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.078 | 0.091 |
| Meta-epidemiology (narrow) | 0.002 | 0.003 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.004 | 0.002 |
| Science and technology studies | 0.005 | 0.004 |
| Scholarly communication | 0.005 | 0.004 |
| Open science | 0.002 | 0.005 |
| Research integrity | 0.005 | 0.007 |
| Insufficient payload (model declined to judge) | 0.103 | 0.051 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".