MétaCan
Menu
Back to cohort

Data extraction methods for systematic review (semi)automation: Update of a living systematic review

2025· preprint· en· W4409247985 on OpenAlexaff
Lena Schmidt, Ailbhe N. Finnerty Mutlu, Rebecca Elmore, Babatunde Kazeem Olorisade, James Thomas, Julian P. T. Higgins

Bibliographic record

VenueF1000Research · 2025
Typepreprint
Languageen
FieldDecision Sciences
TopicMeta-analysis and systematic reviews
Canadian institutionsNortel (Canada)
FundersDepartment of Health and Social CareMedical Research CouncilNational Institute for Health and Care Research
KeywordsOpen peer reviewPlant biologyData extractionMedicinePhysiologyData scienceComputer scienceComputational biologyBiologyMEDLINE

Abstract

fetched live from OpenAlex

Background The reliable and usable (semi) automation of data extraction can support the field of systematic review by reducing the workload required to gather information about the conduct and results of the included studies. This living systematic review examines published approaches for data extraction from reports of clinical studies. Methods We systematically and continually search PubMed, ACL Anthology, arXiv, OpenAlex via EPPI-Reviewer, and the dblp computer science bibliography databases. Full text screening and data extraction are conducted using a mix of open-source and commercial tools. This living review update includes publications up to August 2024 and OpenAlex content up to September 2024. Results 117 publications are included in this review. Of these, 30 (26%) used full texts while the rest used titles and abstracts. A total of 112 (96%) publications developed classifiers for randomised controlled trials. Over 30 entities were extracted, with PICOs (population, intervention, comparator, outcome) being the most frequently extracted. Data are available from 53 (45%), and code from 49 (42%) publications. Nine (8%) implemented publicly available tools. Conclusions This living systematic review presents an overview of (semi)automated data-extraction literature of interest to different types of literature review. We identified a broad evidence base of publications describing data extraction for interventional reviews and a small number of publications extracting other study types. Between review updates, large language models emerged as a new tool for data extraction. While facilitating access to automated extraction, they showed a trend of decreasing quality of results reporting, especially quantitative results such as recall and lower reproducibility of results. Compared with the previous update, trends such as transition to relation extraction and sharing of code and datasets stayed similar.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.365
metaresearch head score (Gemma)0.646
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.635
Threshold uncertainty score0.784

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.3650.646
Meta-epidemiology (narrow)0.0050.005
Meta-epidemiology (broad)0.0160.015
Bibliometrics0.0650.050
Science and technology studies0.0040.006
Scholarly communication0.0130.013
Open science0.0070.014
Research integrity0.0050.005
Insufficient payload (model declined to judge)0.0320.011

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.875
GPT teacher head0.708
Teacher spread0.166 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designNot applicable
DomainMethods
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations22
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueF1000ResearchSame topicMeta-analysis and systematic reviewsFrench-language works237,207