The Safety of Digital Mental Health Interventions: Systematic Review and Recommendations
Bibliographic record
Abstract
BACKGROUND: Evidence suggests that digital mental health interventions (DMHIs) for common mental health conditions are effective. However, digital interventions, such as face-to-face therapies, pose risks to patients. A safe intervention is considered one in which the measured benefits outweigh the identified and mitigated risks. OBJECTIVE: This study aims to review the literature to assess how DMHIs assess safety, what risks are reported, and how they are mitigated in both the research and postmarket phases and building on existing recommendations for assessing, reporting, and mitigating safety in the DMHI and standardizing practice. METHODS: PsycINFO, Embase, and MEDLINE databases were searched for studies that addressed the safety of DMHIs. The inclusion criteria were any study that addressed the safety of a clinical DMHI, even if not as a main outcome, in an adult population, and in English. As the outcome data were mainly qualitative in nature, a meta-analysis was not possible, and qualitative analysis was used to collate the results. Quantitative results were synthesized in the form of tables and percentages. To illustrate the use of a single common safety metric across studies, we calculated odds ratios and CIs, wherever possible. RESULTS: Overall, 23 studies were included in this review. Although many of the included studies assessed safety by actively collecting adverse event (AE) data, over one-third (8/23, 35%) did not assess or collect any safety data. The methods and frequency of safety data collection varied widely, and very few studies have performed formal statistical analyses. The main treatment-related reported AE was symptom deterioration. The main method used to mitigate risk was exclusion of high-risk groups. A secondary web-based search found that 6 DMHIs were available for users or patients to use (postmarket phase), all of which used indications and contraindications to mitigate risk, although there was no evidence of ongoing safety review. CONCLUSIONS: The findings of this review show the need for a standardized classification of AEs, a standardized method for assessing AEs to statically analyze AE data, and evidence-based practices for mitigating risk in DMHIs, both in the research and postmarket phases. This review produced 7 specific, measurable, and achievable recommendations with the potential to have an immediate impact on the field, which were implemented across ongoing and future research. Improving the quality of DMHI safety data will allow meaningful assessment of the safety of DMHIs and confidence in whether the benefits of a new DMHI outweigh its risks. TRIAL REGISTRATION: PROSPERO CRD42022333181; https://www.crd.york.ac.uk/prospero/display_record.php?RecordID=333181.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.004 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".