The impact of digital adherence technologies on treatment outcomes, adherence, and patient-reported outcomes in tuberculosis: a systematic review and meta-analysis
Bibliographic record
Abstract
BACKGROUND: Incomplete tuberculosis (TB) treatment adherence may lead to unsuccessful treatment and relapse. Digital adherence technologies (DATs) may allow more person-centric approaches for supporting treatment adherence. We conducted a systematic review (PROSPERO- CRD42022313166) to evaluate the impact of DATs on adherence, treatment outcomes and patient-reported outcomes in persons treated for TB. METHODS: We searched MEDLINE, Embase, CENTRAL, CINAHL, Web of Science and preprints from Europe PMC, and clinicaltrials.gov for relevant literature from January 2000 to March 2024. We considered experimental or cohort studies reporting quantitative comparisons of adherence, treatment outcomes and patient-reported outcomes between a DAT and the standard of care in each setting. We excluded studies where the technology was used only to log visit attendance or for “routine telephone calls” to patients. Risk of bias was assessed using the Cochrane risk of bias assessment tool and the Newcastle- Ottawa Scale. Pre-specified subgroup analyses considered study design, specific DAT interventions as well as income levels in the countries where studies were conducted. RESULTS: Seventy-six studies (total 86,586 participants) were included evaluating SMS-based interventions (k = 18 studies), feature phone-based interventions (k = 8), medication sleeves with phone calls (branded as “99DOTS,” k = 6), video-observed therapy (VOT; k = 18), smartphone apps (k = 7), digital pillboxes (k = 21), ingestible sensors (k = 1), and interventions combining two DATs (k = 2). Overall, the use of DATs was associated with a modest increase in treatment success in TB disease in both RCTs (OR = 1.14 [0.99, 1.30]; I2 = 57%, k = 34, very low certainty evidence) and observational studies (OR = 1.11 [0.94, 1.30]; I2 = 74%, k = 22, very low certainty evidence). Additionally, DAT use was linked to a significant increase in reporting of adverse events in RCTs (OR = 1.57 [1.25, 1.97]; I2 = 12%, k = 6, moderate certainty) while observational studies showed a similar but non-significant finding (OR = 1.39 [0.93, 2.09]; I2 = 0%, k = 3, moderate certainty). VOT was associated with an increased likelihood of treatment completion in TB infection (OR 4.69 [2.08; 10.55]; I2 = 0%, k = 2, low certainty evidence). VOT also increased frequency of adverse event reporting, as demonstrated in RCTs (OR = 1.9 [1.27; 2.84]; I2 = 0%, k = 3, moderate certainty evidence) and a similar but non-significant effect in observational studies (OR = 1.48 [0.91; 2.42]; I2 = 0%, k = 2, low certainty evidence). Other interventions involving smartphone apps were associated with increased treatment success in TB disease, with a significant effect observed in RCTs (OR 2.17 [1.07; 4.4]; I2 = 20%, k = 3, low certainty evidence) and a non-significant effect in observational studies (OR 1.51 [0.53; 4.3]; I2 = 60%, k = 3, very low certainty evidence). In contrast, interventions with 99DOTS were not associated with improvements in short-term clinical outcomes. There was substantial methodological heterogeneity among studies reporting on adherence. Few studies assessed patient-reported outcomes, though satisfaction was generally higher with DATs. CONCLUSION: Some DATs, notably VOT and smartphone apps, have been successfully used to support TB treatment. Although in many cases DATs did not improve clinical outcomes, they may improve efficiency and adherence, and may be preferred to traditional directly observed therapy by persons with TB. However, evidence remains highly variable, and generalizability limited. Higher quality data are needed. TRIAL REGISTRATION: PROSPERO- CRD42022313166
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.024 | 0.060 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.025 | 0.046 |
| Bibliometrics | 0.009 | 0.010 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.003 | 0.003 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".