MétaCan
Menu
← Back to cohort
Record W4416821455 · doi:10.2196/73492

Evidence for Digital Mental Health Assessment Tools in the Post–COVID-19 Era: Protocol for a Systematic Review on Diagnostic Accuracy Across Age Groups

2025· article· en· W4416821455 on OpenAlexvenueno aff
Kathryn Babbitt, Erin L. Funnell, Nayra A Martin-Key, Eleanor Barker, Sabine Bahn

Bibliographic record

VenueJMIR Research Protocols · 2025
Typearticle
Languageen
FieldPsychology
TopicDigital Mental Health Interventions
Canadian institutionsnot available
Fundersnot available
KeywordsMental healthProtocol (science)Digital healthSoftware deploymentTelemedicineTelehealthHealth careQuality (philosophy)Pandemic

Abstract

fetched live from OpenAlex

Background: Digital assessment tools in health care are increasingly used to aid clinicians in diagnosing mental health conditions. Particularly since the quarantine and isolation guidelines of the COVID-19 pandemic moved much of health care online, there has been an accelerated adoption of digital assessment tools. The diagnostic accuracy of digital mental health assessment tools for a range of psychiatric conditions has yet to be fully explored, especially for their use in populations of older adults and children. Objective: This systematic review aims to (1) summarize recent studies on digital self-report question-and-answer-based mental health assessment tools for use in all ages across a range of psychiatric conditions (eg, the type and number of questions, if available; reference tests; timing; and blinding procedures), (2) present their validity (ie, diagnostic accuracy), and (3) assess study quality and applicability. Methods: The PRISMA-P (Preferred Reporting Items for Systematic Review and Meta-Analysis Protocols) guided the development of this protocol. The protocol has been registered with PROSPERO. The searches were guided by the PICO (population, intervention, comparator, and outcome) framework. A systematic search was conducted of the following databases of literature published since 2021: MEDLINE, Embase, Cochrane Library, ASSIA, Web of Science Core Collection, CINAHL, and PsycINFO. Searches of clinical trial databases and hand searching of reference lists will be completed. Two authors have independently screened titles and abstracts of identified papers and selected studies according to eligibility criteria, resolving inconsistencies through discussion. Full texts were screened following the same process. The authors extracted data using the Covidence data extraction tool (Veritas Health Innovation Ltd; eg, sensitivity and specificity). Two authors will use the Quality Assessment of Diagnostic Accuracy Studies 2 (QUADAS-2) tool to assess risk of bias for each full-text inclusion. Results: Scoping for this review began in December 2024. Searches of databases were completed in January 2025. Full-text screening and identification of the relevant gray literature were completed by the end of August 2025, and the final review is expected to be completed by December 2025. Conclusions: The review aims to present the validity and quality of the diagnostic accuracy of digital mental health assessment tools across different ages (including children and older adults), particularly following the COVID-19 pandemic due to the exponential increase in development and use of such tools. This review will provide evidence for the wider deployment of digital mental health assessment tools across a wide age range. There will also be a discussion about future research for digital tools and avenues for policy around digital mental health assessments.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.065
metaresearch head score (Gemma)0.107
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Systematic review · Consensus signal: Systematic review
GenreCandidate signal: Protocol · Consensus signal: Protocol
Teacher disagreement score0.065
Threshold uncertainty score0.344

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0650.107
Meta-epidemiology (narrow)0.0060.005
Meta-epidemiology (broad)0.0220.024
Bibliometrics0.0170.015
Science and technology studies0.0040.005
Scholarly communication0.0090.011
Open science0.0050.006
Research integrity0.0080.007
Insufficient payload (model declined to judge)0.0560.006

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.560
GPT teacher head0.728
Teacher spread0.168 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSystematic review
Domainnot available
GenreProtocol

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJMIR Research Protocols→Same topicDigital Mental Health Interventions→French-language works237,207→