Comparing Digital Versus Face-to-Face Delivery of Systemic Psychotherapy Interventions: Systematic Review and Meta-Analysis of Randomized Controlled Trials
Bibliographic record
Abstract
BACKGROUND: As digital mental health delivery becomes increasingly prominent, a solid evidence base regarding its efficacy is needed. OBJECTIVE: This study aims to synthesize evidence on the comparative efficacy of systemic psychotherapy interventions provided via digital versus face-to-face delivery modalities. METHODS: We followed PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines for searching PubMed, Embase, Cochrane CENTRAL, CINAHL, PsycINFO, and PSYNDEX and conducting a systematic review and meta-analysis. We included randomized controlled trials comparing mental, behavioral, and somatic outcomes of systemic psychotherapy interventions using self- and therapist-guided digital versus face-to-face delivery modalities. The risk of bias was assessed with the revised Cochrane Risk of Bias tool for randomized trials. Where appropriate, we calculated standardized mean differences and risk ratios. We calculated separate mean differences for nonaggregated analysis. RESULTS: We screened 3633 references and included 12 articles reporting on 4 trials (N=754). Participants were youths with poor diabetic control, traumatic brain injuries, increased risk behavior likelihood, and parents of youths with anorexia nervosa. A total of 56 outcomes were identified. Two trials provided digital intervention delivery via videoconferencing: one via an interactive graphic interface and one via a web-based program. In total, 23% (14/60) of risk of bias judgments were high risk, 42% (25/60) were some concerns, and 35% (21/60) were low risk. Due to heterogeneity in the data, meta-analysis was deemed inappropriate for 96% (54/56) of outcomes, which were interpreted qualitatively instead. Nonaggregated analyses of mean differences and CIs between delivery modalities yielded mixed results, with superiority of the digital delivery modality for 18% (10/56) of outcomes, superiority of the face-to-face delivery modality for 5% (3/56) of outcomes, equivalence between delivery modalities for 2% (1/56) of outcomes, and neither superiority of one modality nor equivalence between modalities for 75% (42/56) of outcomes. Consequently, for most outcome measures, no indication of superiority or equivalence regarding the relative efficacy of either delivery modality can be made at this stage. We further meta-analytically compared digital versus face-to-face delivery modalities for attrition (risk ratio 1.03, 95% CI 0.52-2.03; P=.93) and number of sessions attended (standardized mean difference -0.11; 95% CI -1.13 to -0.91; P=.83), finding no significant differences between modalities, while CIs falling outside the range of the minimal important difference indicate that equivalence cannot be determined at this stage. CONCLUSIONS: Evidence on digital and face-to-face modalities for systemic psychotherapy interventions is largely heterogeneous, limiting conclusions regarding the differential efficacy of digital and face-to-face delivery. Nonaggregated and meta-analytic analyses did not indicate the superiority of either delivery condition. More research is needed to conclude if digital and face-to-face delivery modalities are generally equivalent or if-and in which contexts-one modality is superior to another. TRIAL REGISTRATION: PROSPERO CRD42022335013; https://tinyurl.com/nprder8h.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.040 | 0.108 |
| Meta-epidemiology (narrow) | 0.004 | 0.002 |
| Meta-epidemiology (broad) | 0.029 | 0.042 |
| Bibliometrics | 0.010 | 0.010 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.005 | 0.004 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.003 | 0.003 |
| Insufficient payload (model declined to judge) | 0.005 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".