Clinical Effectiveness, Feasibility, Acceptability, and Usability in Mobile Health Applications for Epilepsy: A Systematic Review (Preprint)
Bibliographic record
Abstract
<sec> <title>BACKGROUND</title> Mobile applications, or “apps”, are widely used by people with epilepsy, their caregivers, and providers. The impact of these apps on the clinical effectiveness (CE) and feasibility, acceptability, or usability (FAU) in epilepsy remains unclear. </sec> <sec> <title>OBJECTIVE</title> To conduct a systematic review of studies investigating the CE and FAU of mobile applications in epilepsy. </sec> <sec> <title>METHODS</title> This review followed the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) standards and was registered with the Prospective Register of Systematic Reviews (PROSPERO; CRD42019134848). The search was conducted using MEDLINE ALL (Ovid) and EMBASE (Ovid) from database inception to April 2022. At the screening phase, we excluded conference abstracts, non-English language and review articles, as well as articles studying video telehealth. We determined study quality for case-control or cohort studies using the Newcastle-Ottawa Quality Assessment Scale (NOQAS) and bias in randomized studies using the Cochrane Collaboration Handbook Risk of Bias (RoB) tool. We assessed usability study quality using the validated 15-point Silva scale. Study characteristics were analyzed using summary statistics. </sec> <sec> <title>RESULTS</title> We identified 6,768 studies, of which 13 (0.2%) were included. Of the 13 studies, 8 (61.5%) addressed CE, 6 (46.2%) acceptability, 5 (38.5%) usability, and 4 (30.8%) feasibility. Four studies (31.0%) evaluated both CE and FAU. Studies comprised prospective cohort (N=6, 46.2%), pilot (N=3, 23.1%), randomized trial (N=3, 23.1%) and pre/post (N=1, 7.7%) designs. Overall, cohort studies demonstrated fair quality (median NOQAS score 5, interquartile range [IQR] 5.0 - 5.8), whereas 2 (66.7%) randomized studies had some concern for bias. Usability studies demonstrated high methodological quality (median Silva score 10, IQR 10 - 11). Apps were most frequently studied in patient users (N=7 (87.5%) CE and 8 (100%) FAU studies). The most common app target in CE studies was physical health (N=5, 62.5%) contrasting with symptom management (N=7, 87.5%) in FAU studies. </sec> <sec> <title>CONCLUSIONS</title> We found that studies of app use in epilepsy most commonly studied CE and evaluated patient-facing apps. Despite high methodological quality in usability studies and several randomized CE studies, cohort and randomized studies demonstrated fair quality and moderate bias, respectively. Additional high-quality evidence is necessary to evaluate the CE and FAU of app use in epilepsy. </sec>
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | Meta-epidemiology (broad) Domain: not available · Genre: Review About the Canadian research system: no · About a Canadian topic: no | Systematic review | low |
| gpt | Meta-epidemiology (broad) Domain: not available · Genre: Review About the Canadian research system: no · About a Canadian topic: no | Systematic review | high |
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.052 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.011 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".