Female authorship trends in a high-impact Canadian medical journal: a 10-year cross-sectional series, 2013–2023
Bibliographic record
Abstract
IMPORTANCE: Women are under-represented in senior roles within academic medicine, including as authors in high-impact journals. OBJECTIVE: To examine trends and predictors of female authorship in the Canadian Medical Association Journal (CMAJ) as the only high-impact Canadian journal over a 10-year period to understand gender balances in Canadian academic publishing. DESIGN: This cross-sectional study analysed trends and predictors of female authorship in articles published in CMAJ from 1 January 2013 to 31 December 2023. SETTING: Data were extracted from PubMed for CMAJ, the only high-impact Canadian medical journal (impact factor ≥10). Data extraction used the RISmed package in R Studio. PARTICIPANTS: The study included articles published in CMAJ within the specified period. Author gender was predicted using the validated Genderize.io software. Articles where the gender of the authors could not be predicted were excluded from analysis. MAIN OUTCOMES AND MEASURES: tests comparing proportions, Jonckheere and linear regression models to evaluate trends. Among multiauthor articles, multivariable logistic regression models assessed predictors of female first and last authorship. RESULTS: From 5805 included articles, women comprised 47% of first authors and 43% of last authors (p<0.001), both significantly lower than men (p<0.001). Female first authorship increased by 17.7% and female last authorship by 10.5% over the study period (both p<0.05 for trend), reaching a majority (58%) and near parity (48%) in 2023, respectively. Female editor-in-chief and higher proportion of female coauthors were associated with higher odds of female first and last authors; female last authors were additionally associated with higher odds of female first authors. INTERPRETATION: Women were under-represented in authorship overall, though female first and last authorship increased over time, with first authorship exceeding parity in recent years and last authorship nearing equal representation. Female editors-in-chief and a higher proportion of female coauthors were associated with greater female first and last authorship, while female last authorship was additionally associated with higher odds of female first authorship. These findings provide insight into authorship trends in a high-impact Canadian medical journal and may inform future efforts to support gender equity in academic publishing.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.016 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.010 | 0.015 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".