MétaCan
Menu
Back to cohort
Record W4403824124 · doi:10.1093/eurpub/ckae144.1053

Using artificial intelligence to support and streamline rapid systematic evidence reviews

2024· article· en· W4403824124 on OpenAlexaff
Michael A. Dobbins, Robyn Traynor, Emily Clark, Sarah Neil‐Sztramko

Bibliographic record

VenueEuropean Journal of Public Health · 2024
Typearticle
Languageen
FieldDecision Sciences
TopicMeta-analysis and systematic reviews
Canadian institutionsImpactMcMaster University
Fundersnot available
KeywordsSystematic reviewComputer scienceData scienceMEDLINEChemistry

Abstract

fetched live from OpenAlex

Abstract Issue The Rapid Evidence Service, initiated by the National Collaborating Centre for Methods and Tools (NCCMT) during COVID-19, supports public health decision making by conducting rapid reviews on priority topics. Description of issue Integral to rapid reviews is an expedited timeline but the quantity of available literature for most public health review questions takes significant time to screen manually. NCCMT integrated 4 artificial intelligence (AI) features into the screening process. DAISY Rank applies predictions learned from manual screening patterns to re-order remaining studies, with most relevant appearing first. AI Screening automatically screens remaining studies based on prediction scores. Check for Screening Errors and Re-Rank Report use previous screening patterns to identify studies that were potentially falsely excluded and predict the total number of included studies, respectively. These features were tested by comparing results provided by AI with those produced manually for select test sets. Results NCCMT used AI to support and expedite screening, assess screening progress, and/or minimize risk of inappropriately excluding studies for 35 rapid reviews on 20 topics. Using DAISY Rank enabled one screener to review over 4000 references in 9 hours, compared to a different review, where the same amount of screening took 28 hours without DAISY Rank. AI Screening correctly excluded up to 80% of irrelevant search results across reviews. Check for Screening Errors identified 37 potential includes manually excluded in one review; these were reviewed and 3 were included. Re-Rank Report allowed NCCMT to re-allocate staff to subsequent steps in the review process when most included studies were identified. Lessons Integrating AI features into screening led to less time required, better anticipated timelines, more accurate staff allocation and reduced errors. More rigorous study of AI best practices is needed to continue to improve rapid review method efficiencies. Key messages • Rapid reviews can be an important source of evidence for decision makers if they can be completed quickly but maintain rigor and accuracy. • AI holds promise as a way to improve screening efficiency.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.278
metaresearch head score (Gemma)0.648
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.722
Threshold uncertainty score0.890

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.2780.648
Meta-epidemiology (narrow)0.0030.004
Meta-epidemiology (broad)0.0050.005
Bibliometrics0.0260.015
Science and technology studies0.0020.001
Scholarly communication0.0140.010
Open science0.0050.008
Research integrity0.0030.004
Insufficient payload (model declined to judge)0.0060.004

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.941
GPT teacher head0.607
Teacher spread0.334 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designTheoretical or conceptual
DomainMethods
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueEuropean Journal of Public HealthSame topicMeta-analysis and systematic reviewsFrench-language works237,207