MétaCan
Menu
Back to cohort
Record W4378190436 · doi:10.1002/jrsm.1636

A real‐world evaluation of the implementation of <scp>NLP</scp> technology in abstract screening of a systematic review

2023· review· en· W4378190436 on OpenAlexafffund
Sara Perlman‐Arrow, Noel Loo, Niklas Bobrovitz, Tingting Yan, Rahul K. Arora

Bibliographic record

VenueResearch Synthesis Methods · 2023
Typereview
Languageen
FieldDecision Sciences
TopicMeta-analysis and systematic reviews
Canadian institutionsUniversity of TorontoMcGill UniversityThe Quebec Population Health Research NetworkUniversity of CalgaryMcGill University Health Centre
FundersCanadian Medical AssociationRobert Koch InstitutKoch Institute for Integrative Cancer Research, Massachusetts Institute of TechnologyPublic Health AgencyRhodes ScholarshipsPublic Health Agency of CanadaWorld Health Organization
KeywordsComputer scienceContext (archaeology)Systematic reviewRecallInclusion (mineral)Machine learningArtificial intelligencePrecision and recallNatural language processingMEDLINEPsychology

Abstract

fetched live from OpenAlex

The laborious and time-consuming nature of systematic review production hinders the dissemination of up-to-date evidence synthesis. Well-performing natural language processing (NLP) tools for systematic reviews have been developed, showing promise to improve efficiency. However, the feasibility and value of these technologies have not been comprehensively demonstrated in a real-world review. We developed an NLP-assisted abstract screening tool that provides text inclusion recommendations, keyword highlights, and visual context cues. We evaluated this tool in a living systematic review on SARS-CoV-2 seroprevalence, conducting a quality improvement assessment of screening with and without the tool. We evaluated changes to abstract screening speed, screening accuracy, characteristics of included texts, and user satisfaction. The tool improved efficiency, reducing screening time per abstract by 45.9% and decreasing inter-reviewer conflict rates. The tool conserved precision of article inclusion (positive predictive value; 0.92 with tool vs. 0.88 without) and recall (sensitivity; 0.90 vs. 0.81). The summary statistics of included studies were similar with and without the tool. Users were satisfied with the tool (mean satisfaction score of 4.2/5). We evaluated an abstract screening process where one human reviewer was replaced with the tool's votes, finding that this maintained recall (0.92 one-person, one-tool vs. 0.90 two tool-assisted humans) and precision (0.91 vs. 0.92) while reducing screening time by 70%. Implementing an NLP tool in this living systematic review improved efficiency, maintained accuracy, and was well-received by researchers, demonstrating the real-world effectiveness of NLP in expediting evidence synthesis.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.665
metaresearch head score (Gemma)0.825
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: Methods
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.335
Threshold uncertainty score0.413

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.6650.825
Meta-epidemiology (narrow)0.0030.003
Meta-epidemiology (broad)0.0040.010
Bibliometrics0.0110.013
Science and technology studies0.0030.004
Scholarly communication0.0100.008
Open science0.0040.006
Research integrity0.0050.003
Insufficient payload (model declined to judge)0.0050.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.951
GPT teacher head0.759
Teacher spread0.192 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designObservational
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations27
Published2023
Admission routes2
Has abstractyes

Explore more

Same venueResearch Synthesis MethodsSame topicMeta-analysis and systematic reviewsFrench-language works237,207