How Notifications Affect Engagement With a Behavior Change App: Results From a Micro-Randomized Trial
Bibliographic record
Abstract
BACKGROUND: Drink Less is a behavior change app to help higher-risk drinkers in the United Kingdom reduce their alcohol consumption. The app includes a daily notification asking users to "Please complete your drinks and mood diary," yet we did not understand the causal effect of the notification on engagement nor how to improve this component of Drink Less. We developed a new bank of 30 new messages to increase users' reflective motivation to engage with Drink Less. This study aimed to determine how standard and new notifications affect engagement. OBJECTIVE: Our objective was to estimate the causal effect of the notification on near-term engagement, to explore whether this effect changed over time, and to create an evidence base to further inform the optimization of the notification policy. METHODS: We conducted a micro-randomized trial (MRT) with 2 additional parallel arms. Inclusion criteria were Drink Less users who consented to participate in the trial, self-reported a baseline Alcohol Use Disorders Identification Test score of ≥8, resided in the United Kingdom, were aged ≥18 years, and reported interest in drinking less alcohol. Our MRT randomized 350 new users to test whether receiving a notification, compared with receiving no notification, increased the probability of opening the app in the subsequent hour, over the first 30 days since downloading Drink Less. Each day at 8 PM, users were randomized with a 30% probability of receiving the standard message, a 30% probability of receiving a new message, or a 40% probability of receiving no message. We additionally explored time to disengagement, with the allocation of 60% of eligible users randomized to the MRT (n=350) and 40% of eligible users randomized in equal number to the 2 parallel arms, either receiving the no notification policy (n=98) or the standard notification policy (n=121). Ancillary analyses explored effect moderation by recent states of habituation and engagement. RESULTS: Receiving a notification, compared with not receiving a notification, increased the probability of opening the app in the next hour by 3.5-fold (95% CI 2.91-4.25). Both types of messages were similarly effective. The effect of the notification did not change significantly over time. A user being in a state of already engaged lowered the new notification effect by 0.80 (95% CI 0.55-1.16), although not significantly. Across the 3 arms, time to disengagement was not significantly different. CONCLUSIONS: We found a strong near-term effect of engagement on the notification, but no overall difference in time to disengagement between users receiving the standard fixed notification, no notification at all, or the random sequence of notifications within the MRT. The strong near-term effect of the notification presents an opportunity to target notifications to increase "in-the-moment" engagement. Further optimization is required to improve the long-term engagement. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): RR2-10.2196/18690.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.037 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.006 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.003 | 0.004 |
| Insufficient payload (model declined to judge) | 0.016 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".