Scholarly publishing’s hidden diversity: How exclusive databases sustain the oligopoly of academic publishers
Bibliographic record
Abstract
Global scholarly publishing has been dominated by a small number of publishers for several decades. This paper revisits the data on corporate control of scholarly publishing by analyzing the relative shares of scholarly journals and articles published by the major publishers and the "long tail" of smaller, independent publishers, using Dimensions and Web of Science (WoS). The reduction of expenses for printing and distribution and the availability of open-source journal management tools may have contributed to the emergence of small publishers, while recently developed inclusive databases may allow for the study of these. Dimensions' inclusive indexing revealed the number of scholarly journals and articles published by smaller publishers has been growing rapidly, especially since the onset of large-scale online publishing around 2000, resulting in a higher share of articles from smaller publishers. In parallel, WoS shows increasing concentration within a few corporate publishers. For the 1980-2021 period, we retrieved 32% more articles from Dimensions compared to the more selective WoS. Dimensions' data showed the expansion of small publishers was most pronounced in the Social Sciences and the Arts and Humanities, but a similar trend is observed in the Natural Sciences and Engineering, and the Health Sciences. A major geographical divergence is also revealed, with English-speaking countries and/or those located in northwestern Europe relying heavily on major publishers for the dissemination of their research, while the rest of the world being relatively independent of the oligopoly. Finally, independent journals publish more often in open access in general, and in Diamond open access in particular. We conclude that enhanced indexing and visibility of recently created, independent journals may favour their growth and stimulate global scholarly bibliodiversity.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.019 | 0.128 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.008 | 0.014 |
| Science and technology studies | 0.005 | 0.009 |
| Scholarly communication | 0.027 | 0.031 |
| Open science | 0.003 | 0.012 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.008 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".