MétaCan
Menu
Back to cohort
Record W4406712574 · doi:10.1136/bmjopen-2024-086982

Diversity in the medical research ecosystem: a descriptive scientometric analysis of over 49 000 studies and 150 000 authors published in high-impact medical journals between 2007 and 2022

2025· article· en· W4406712574 on OpenAlexaff
Marie‐Laure Charpignon, João Matos, Luis Filipe Nakayama, Jack Gallifant, Pia Gabrielle I. Alfonso, Marisa Cobanaj, Amelia Fiske, Alexander J. Gates, Frances Dominique V. Ho, Urvish Jain, Mohammad Kashkooli, Naira Link, Liam G. McCoy, Jonathan D. Shaffer, Leo Anthony Celi

Bibliographic record

VenueBMJ Open · 2025
Typearticle
Languageen
FieldSocial Sciences
TopicDiversity and Career in Medicine
Canadian institutionsUniversity of Alberta
FundersNational Institute of Biomedical Imaging and BioengineeringFulbright PortugalFundação Luso-Americana para o DesenvolvimentoBroad Institute
KeywordsMedicineDiversity (politics)Impact factorScale (ratio)Medical journalBibliometricsDescriptive statisticsRepresentation (politics)Family medicineLibrary scienceGeographyPolitical science

Abstract

fetched live from OpenAlex

Objectives Health research that significantly impacts global clinical practice and policy is often published in high-impact factor (IF) medical journals. These outlets play a pivotal role in the worldwide dissemination of novel medical knowledge. However, researchers identifying as women and those affiliated with institutions in low- and middle-income countries (LMICs) have been largely under-represented in high-IF journals across multiple fields of medicine. To evaluate disparities in gender and geographical representation among authors who have published in any of five top general medical journals, we conducted scientometric analyses using a large-scale dataset extracted from the New England Journal of Medicine , Journal of the American Medical Association , The BMJ , The Lancet and Nature Medicine . Methods Author metadata from all articles published in the selected journals between 2007 and 2022 were collected using the DimensionsAI platform. The Genderize.io Application Programming Interface was then used to infer each author’s likely gender based on their extracted first name. The World Bank country classification was used to map countries associated with researcher affiliations to the LMIC or the high-income country (HIC) category. We characterised the overall gender and country income category representation across the five medical journals. In addition, we computed article-level diversity metrics and contrasted their distributions across the journals. Results We studied 151 536 authors across 49 764 articles published in five top medical journals, over a period spanning 15 years. On average, approximately one-third (33.1%) of the authors of a given paper were inferred to be women; this result was consistent across the journals we studied. Further, 86.6% of the teams were exclusively composed of HIC authors; in contrast, only 3.9% were exclusively composed of LMIC authors. The probability of serving as the first or last author was significantly higher if the author was inferred to be a man (18.1% vs 16.8%, p<0.01) or was affiliated with an institution in a HIC (16.9% vs 15.5%, p<0.01). Our primary finding reveals that having a diverse team promotes further diversity, within the same dimension (ie, gender or geography) and across dimensions. Notably, papers with at least one woman among the authors were more likely to also involve at least two LMIC authors (11.7% vs 10.4% in baseline, p<0.001; based on inferred gender); conversely, papers with at least one LMIC author were more likely to also involve at least two women (49.4% vs 37.6%, p<0.001; based on inferred gender). Conclusion We provide a scientometric framework to assess authorship diversity. Our research suggests that the inclusiveness of high-impact medical journals is limited in terms of both gender and geography. We advocate for medical journals to adopt policies and practices that promote greater diversity and collaborative research. In addition, our findings offer a first step towards understanding the composition of teams conducting medical research globally and an opportunity for individual authors to reflect on their own collaborative research practices and possibilities to cultivate more diverse partnerships in their work.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Direct model labels (unvalidated)

Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.

Model armCategoriesStudy designConfidence
gemmaMetaresearchBibliometrics
Domain: Incentives · Genre: Empirical
About the Canadian research system: no · About a Canadian topic: no
Observationallow
gptBibliometricsMetaresearch
Domain: Incentives · Genre: Empirical
About the Canadian research system: no · About a Canadian topic: no
Other designhigh
models splitAgreement compares identical category sets and study designs across arms.

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.099
metaresearch head score (Gemma)0.022
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Insufficient payload (model declined to judge)
Consensus categoriesMetaresearch
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.077
Threshold uncertainty score0.999

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0990.022
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0040.017
Science and technology studies0.0010.001
Scholarly communication0.0000.001
Open science0.0020.004
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0020.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.323
GPT teacher head0.554
Teacher spread0.231 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Labeled directly by 2 models reading the full record.

MetaresearchBibliometrics

The models disagree on parts of this classification; every voice is preserved in the section at the end of the page.

Study designObservational · Other design
DomainIncentives
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations8
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueBMJ OpenSame topicDiversity and Career in MedicineCategoryMetaresearchFrench-language works237,207