MétaCan
Menu
Back to cohort
Record W4200362342 · doi:10.1186/s12874-021-01451-2

Guidance for using artificial intelligence for title and abstract screening while conducting knowledge syntheses

2021· article· en· W4200362342 on OpenAlexafffund
Candyce Hamel, Mona Hersi, Shannon Kelly, Andrea C. Tricco, Sharon E. Straus, George A. Wells, Ba’ Pham, Brian Hutton

Bibliographic record

VenueBMC Medical Research Methodology · 2021
Typearticle
Languageen
FieldMedicine
TopicArtificial Intelligence in Healthcare and Education
Canadian institutionsUniversity of TorontoSt. Michael's HospitalUniversity of OttawaOttawa Hospital
FundersCanadian Institutes of Health Research
KeywordsWorkflowComputer scienceCornerstoneSet (abstract data type)Systematic reviewCitationArtificial intelligenceApplications of artificial intelligenceProcess (computing)Knowledge managementData scienceMEDLINEDatabaseWorld Wide Web

Abstract

fetched live from OpenAlex

BACKGROUND: Systematic reviews are the cornerstone of evidence-based medicine. However, systematic reviews are time consuming and there is growing demand to produce evidence more quickly, while maintaining robust methods. In recent years, artificial intelligence and active-machine learning (AML) have been implemented into several SR software applications. As some of the barriers to adoption of new technologies are the challenges in set-up and how best to use these technologies, we have provided different situations and considerations for knowledge synthesis teams to consider when using artificial intelligence and AML for title and abstract screening. METHODS: We retrospectively evaluated the implementation and performance of AML across a set of ten historically completed systematic reviews. Based upon the findings from this work and in consideration of the barriers we have encountered and navigated during the past 24 months in using these tools prospectively in our research, we discussed and developed a series of practical recommendations for research teams to consider in seeking to implement AML tools for citation screening into their workflow. RESULTS: We developed a seven-step framework and provide guidance for when and how to integrate artificial intelligence and AML into the title and abstract screening process. Steps include: (1) Consulting with Knowledge user/Expert Panel; (2) Developing the search strategy; (3) Preparing your review team; (4) Preparing your database; (5) Building the initial training set; (6) Ongoing screening; and (7) Truncating screening. During Step 6 and/or 7, you may also choose to optimize your team, by shifting some members to other review stages (e.g., full-text screening, data extraction). CONCLUSION: Artificial intelligence and, more specifically, AML are well-developed tools for title and abstract screening and can be integrated into the screening process in several ways. Regardless of the method chosen, transparent reporting of these methods is critical for future studies evaluating artificial intelligence and AML.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.642
metaresearch head score (Gemma)0.847
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: Methods
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.358
Threshold uncertainty score0.442

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.6420.847
Meta-epidemiology (narrow)0.0050.006
Meta-epidemiology (broad)0.0060.012
Bibliometrics0.0480.029
Science and technology studies0.0070.009
Scholarly communication0.0300.030
Open science0.0140.018
Research integrity0.0120.016
Insufficient payload (model declined to judge)0.0160.020

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.966
GPT teacher head0.707
Teacher spread0.259 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designNot applicable
DomainMethods
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations101
Published2021
Admission routes2
Has abstractyes

Explore more

Same venueBMC Medical Research MethodologySame topicArtificial Intelligence in Healthcare and EducationFrench-language works237,207