Developing a Behavioral Phenotyping Layer for Artificial Intelligence–Driven Predictive Analytics in a Digital Resiliency Course: Protocol for a Randomized Controlled Trial
Bibliographic record
Abstract
BACKGROUND: Digital interventions for mental health are pivotal for addressing barriers such as stigma, cost, and accessibility, particularly for underserved populations. While the effectiveness of digital interventions has been established, poor adherence and lack of engagement remain critical factors that undermine efficacy. Millions of individuals will never have access to a trained mental health care practitioner, underscoring the need for highly tailored and engaging self-guided resources. This study builds on a prior study that successfully leveraged behavioral economics (nudges and prompts) to enhance engagement. Expanding on that study, this research will focus on building a foundational dataset of behavioral phenotypes to support artificial intelligence (AI)-driven personalization in digital mental health. OBJECTIVE: This 6-arm randomized controlled trial aims to analyze user engagement with randomized tips and to-do lists within a resiliency course tailored for Ukrainian refugees affected by the ongoing humanitarian crisis (Спільна Сила), using the EvolutionHealth.care (V-CC Systems Inc) platform. Insights will inform the development of an AI-based personalization system to optimize engagement and address behavioral health challenges. Secondary objectives include identifying demographic and behavioral predictors of engagement and creating a scalable, culturally sensitive intervention model. METHODS: Participants will be recruited through digital outreach, enrolled anonymously, and randomized into 6 groups to compare combinations of tips, nudges, and to-do lists. Engagement metrics (eg, clicks, completion rates, and session duration) and demographic data (eg, age and gender) will be collected. Statistical analyses will include a comparison between arms and interaction testing to evaluate the effectiveness of each intervention component. Ethical safeguards include institutional review board approval, informed consent, and strict data privacy standards. RESULTS: This protocol was designed in January 2025. α and β testing of the intervention are scheduled to begin in July 2025, with a soft launch anticipated in August 2025. The experiment will remain active until the sample size requirements are met. Live monitoring and periodic data quality checks will be conducted throughout the study duration. CONCLUSIONS: This trial represents a novel approach to behavioral health research by leveraging randomized experimentation to develop AI-ready behavioral datasets. By targeting an underserved and culturally sensitive population, it contributes critical insights toward scalable, personalized digital mental health interventions. Findings may help inform future digital health efforts that aim to improve engagement, accessibility, and long-term adherence. TRIAL REGISTRATION: Open Science Framework 34rmg; https://osf.io/34rmg. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): PRR1-10.2196/73773.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.051 | 0.065 |
| Meta-epidemiology (narrow) | 0.006 | 0.003 |
| Meta-epidemiology (broad) | 0.009 | 0.006 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.004 | 0.004 |
| Scholarly communication | 0.005 | 0.005 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.007 | 0.010 |
| Insufficient payload (model declined to judge) | 0.106 | 0.019 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".