MétaCan
Menu
Back to cohort
Record W4411259709 · doi:10.1101/2025.06.13.25329541

Automation of Systematic Reviews with Large Language Models

2025· preprint· en· W4411259709 on OpenAlexaffabout
Christian Cao, Rohit Arora, Paul Cento, Katherine Manta, Elina Farahani, Milena Cecere, Anabel Selemon, Jason C. Sang, L. Gong, Robert Kloosterman, Richard Saleh, Denis A. Margalik, Lin James, Jane Jomy, David Chen, Jaswanth Gorla, S.-W. Lee, Kelvin Zhang, Mairead Whelan, Bijan Teja, Alexander A. C. Leung, Rahul K. Arora, Michael Noetel, Niklas Bobrovitz

Bibliographic record

VenuemedRxiv · 2025
Typepreprint
Languageen
FieldComputer Science
TopicAdvanced Text Analysis Techniques
Canadian institutionsPublic Health OntarioUniversity Health NetworkOttawa HospitalUniversity of AlbertaSt. Michael's HospitalUniversity of British ColumbiaMount Sinai HospitalUniversity of CalgaryMcGill UniversityUniversity of OttawaVector InstituteWilfrid Laurier UniversityUniversity of WaterlooUniversity of Toronto
Fundersnot available
KeywordsAutomationComputer scienceSystems engineeringEngineeringMechanical engineering

Abstract

fetched live from OpenAlex

Abstract Importance Systematic reviews (SRs) inform evidence-based decision making. Yet, many take over a year to complete, are labor intensive, prone to human error, and face reproducibility challenges; thus limiting access to timely and reliable information. Objective To validate a large language model (LLM)-based workflow (otto-SR) to automate three of the most labour intensive tasks in performing SR’s: article screening, data extraction, and risk of bias assessment; and to assess its feasibility in rapidly updating existing reviews. Design, setting, and participants We conducted a validation study in four phases, with direct benchmarking against graduate-level human researchers in phases 1 and 2. Phase 1: article screening performance was measured across 32,357 citations from 5 systematic reviews. The reference standard consisted of the original reviews’ screening decisions after full-text screening. Phase 2: data extraction performance was measured across 4,495 data points from 495 studies in 7 reviews. Phase 3: risk of bias assessment (ROB2, Newcastle-Ottawa, QUADAS2) performance was measured across 345 studies from 12 reviews. Reference standards for Phase 2 and Phase 3 were created after blinded adjudication of the original review extraction and RoB assessments. Phase 4: otto-SR was used to reproduce and update the primary analysis from an issue of Cochrane reviews (n=12 reviews, 146,276 citations), with analytical comparisons to the original meta-analyzed findings. All discrepancies underwent dual human review. Results otto-SR showed high performance in phase 1 article screening ( otto-SR : 96.7% sensitivity, 97.9% specificity; human: 81.7% sensitivity, 98.1% specificity) and phase 2 data extraction ( otto-SR : 93.1% accuracy; human: 79.7% accuracy). In phase 3, otto-SR demonstrated high interrater reliability for risk of bias judgements (ROB2 0.98, Newcastle-Ottawa 0.95, QUADAS2 0.74; Gwet AC2). In phase 4, otto-SR , reproduced and updated the primary analysis from an issue of Cochrane reviews. Across Cochrane reviews, otto-SR incorrectly excluded a median of 0 studies (IQR 0 to 0.25), and found nearly twice as many eligible studies compared to the original authors (n= 114 vs. 64). Meta-analyses based on otto-SR generated screening and extraction outputs, subsequently verified through dual human review, yielded newly statistically significant effect estimates in 2 reviews and negated significance in 1 review. Conclusions and relevance LLMs have high performance in article screening, data extraction, and risk of bias assessments. They can rapidly reproduce and update existing systematic reviews, laying the foundation for automated, scalable, and reliable evidence synthesis.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.345
metaresearch head score (Gemma)0.690
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.655
Threshold uncertainty score0.807

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.3450.690
Meta-epidemiology (narrow)0.0040.005
Meta-epidemiology (broad)0.0060.010
Bibliometrics0.0150.011
Science and technology studies0.0030.002
Scholarly communication0.0130.009
Open science0.0060.015
Research integrity0.0020.004
Insufficient payload (model declined to judge)0.0090.006

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.034
GPT teacher head0.323
Teacher spread0.289 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designSimulation or modeling
DomainMethods
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations21
Published2025
Admission routes2
Has abstractyes

Explore more

Same venuemedRxivSame topicAdvanced Text Analysis TechniquesFrench-language works237,207