MétaCan
Menu
Back to cohort
Record W4416714491 · doi:10.1016/j.artmed.2025.103312

Do machine learning methods make better predictions than conventional ones in pharmacoepidemiology? A systematic review, meta-analysis, and network meta-analysis

2025· article· en· W4416714491 on OpenAlexafffund
Ana Paula Bruno Pena-Gralle, Mireille E. Schnitzer, Sofia-Nada Boureguaa, F. Morin, Marc‐André Legault, Caroline Sirois, Alice Dragomir, Lucie Blais

Bibliographic record

VenueArtificial Intelligence in Medicine · 2025
Typearticle
Languageen
FieldDecision Sciences
TopicMeta-analysis and systematic reviews
Canadian institutionsUniversité LavalMcGill UniversityHôpital du Sacré-Cœur de MontréalUniversité de Montréal
FundersFonds de Recherche du Québec - SantéCanadian Institutes of Health ResearchRéseau Québécois de Recherche sur les Médicaments
KeywordsArtificial neural networkDeep learningSupport vector machineComputational learning theoryKey (lock)

Abstract

fetched live from OpenAlex

OBJECTIVE: To synthesize existing evidence and compare the predictive performance of conventional statistical (CS) models versus machine learning (ML) methods in pharmacoepidemiology. METHODS: Medline, Embase, PsycINFO, CINAHL and Web of Science databases were systematically searched for predictive pharmacoepidemiologic studies published between January 2018 and September 2025. Independent reviewers extracted predictive metrics and other data from each study and assessed the quality of the comparison between methods. The relative performance of ML compared to CS was estimated for each prediction objective. Performance metrics were pooled in meta-analyses and Bayesian network meta-analyses (NMA). RESULTS: Among 9106 records identified, 65 studies met inclusion criteria, encompassing 83 prediction objectives. For 84 % of these objectives, CS was outperformed by at least one ML method. The median sample size across these studies was 2691 subjects (50 to 1,807,159), and, for binary outcomes, the median number of events per candidate predictor was 17.9 (range 0.28 to 24,260). We observed a risk of bias in the comparison according to at least one of eight major criteria for 39 prediction objectives (47 %). The pooled area under the receiver-operator curve (AUC) ratio for the highest-performing ML method in studies with low risk of bias was estimated as 1.07 (95 % confidence interval 1.03-1.12) in favor of ML, but with very high heterogeneity. NMA of 197 comparisons estimated an AUC ratio of 1.07 (95 % credible interval 1.04-1.12) for boosted methods compared with logistic regression and ranked Gradient Boosting Machine and XGBoost consistently among the best-performing methods. CONCLUSION: Machine learning methods applied to structured pharmacoepidemiologic data demonstrated a consistent yet modest advantage in discriminative performance relative to conventional statistical models. This advantage was most evident for boosted methods such as GBM and XGBoost. However, greater rigor in reporting methodological details is recommended to improve the comprehension, transparency, and reproducibility of studies. REGISTRATION: PROSPERO 2023 registration number: CRD42023426986.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.327
metaresearch head score (Gemma)0.083
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Meta-epidemiology (narrow), Meta-epidemiology (broad), Insufficient payload (model declined to judge)
Consensus categoriesMetaresearch, Meta-epidemiology (broad)
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Meta-analysis · Consensus signal: Meta-analysis
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.862
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.3270.083
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0270.011
Bibliometrics0.0040.021
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0020.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0260.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.769
GPT teacher head0.607
Teacher spread0.162 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designMeta-analysis
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2025
Admission routes2
Has abstractyes

Explore more

Same venueArtificial Intelligence in MedicineSame topicMeta-analysis and systematic reviewsFrench-language works237,207