Understanding Design Approaches and Evaluation Methods in mHealth Apps Targeting Substance Use: Protocol for a Systematic Review
Bibliographic record
Abstract
BACKGROUND: Substance use and use disorders in the United States have had significant and devastating impacts on individuals and communities. This escalating substance use crisis calls for urgent and innovative solutions to effectively detect and provide interventions for individuals in times of need. Recent mobile health (mHealth)-based approaches offer promising new opportunities to address these issues through ubiquitous devices. However, the design rationales, theoretical frameworks, and mechanisms through which users' perspectives and experiences guide the design and deployment of such systems have not been analyzed in any prior systematic reviews. OBJECTIVE: In this paper, we systematically review these approaches and apps for their feasibility, efficacy, and usability. Further, we evaluate whether human-centered research principles and techniques guide the design and development of these systems and examine how the current state-of-the-art systems apply to real-world contexts. In an effort to gauge the applicability of these systems, we also investigate whether these approaches consider the effects of stigma and privacy concerns related to collecting data on substance use. Lastly, we examine persistent challenges in the design and large-scale adoption of substance use intervention apps and draw inspiration from other domains of mHealth to suggest actionable reforms for the design and deployment of these apps. METHODS: Four databases (PubMed, IEEE Xplore, JMIR, and ACM Digital Library) were searched over a 5-year period (2016-2021) for articles evaluating mHealth approaches for substance use (alcohol use, marijuana use, opioid use, tobacco use, and substance co-use). Articles that will be included describe an mHealth detection or intervention targeting substance use, provide outcomes data, and include a discussion of design techniques and user perspectives. Independent evaluation will be conducted by one author, followed by secondary reviewer(s) who will check and validate themes and data. RESULTS: This is a protocol for a systematic review; therefore, results are not yet available. We are currently in the process of selecting the studies for inclusion in the final analysis. CONCLUSIONS: To the best of our knowledge, this is the first systematic review to assess real-world applicability, scalability, and use of human-centered design and evaluation techniques in mHealth approaches targeting substance use. This study is expected to identify gaps and opportunities in current approaches used to develop and assess mHealth technologies for substance use detection and intervention. Further, this review also aims to highlight various design processes and components that result in engaging, usable, and effective systems for substance use, informing and motivating the future development of such systems. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): DERR1-10.2196/35749.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.179 | 0.213 |
| Meta-epidemiology (narrow) | 0.008 | 0.006 |
| Meta-epidemiology (broad) | 0.020 | 0.021 |
| Bibliometrics | 0.023 | 0.020 |
| Science and technology studies | 0.007 | 0.007 |
| Scholarly communication | 0.011 | 0.010 |
| Open science | 0.006 | 0.008 |
| Research integrity | 0.010 | 0.009 |
| Insufficient payload (model declined to judge) | 0.040 | 0.007 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".