Preprint servers and journals: rivals or allies?
Bibliographic record
Abstract
Purpose This study explores the evolving role of preprint servers within the scholarly communication system, focusing on their relationship with peer-reviewed journals. As preprints become more common, questioning and understanding their future role is critical for maintaining a healthy scholarly communication ecosystem. By examining the values, concerns and goals of preprint server managers, this study highlights the significant influence these individuals have in shaping the future of preprints. Design/methodology/approach A qualitative, interview-based approach was used to gather insights from preprint server managers on their roles, challenges and visions for the future of preprints within the broader scholarly communication system. Findings The findings point to a lack of consensus on how preprint servers and journals should interact and to diverging views on how the certification and curation functions are best performed and by whom. Concerns about credibility and long-term financial sustainability are increasingly driving independent and community-run preprint servers to align more closely with journals, potentially undermining the disruptive and emancipatory potential of preprints. Originality/value This study is the first to examine the relationship between preprints and journals from the perspective of preprint server managers in the later stages of the COVID-19 pandemic. It sheds light on how preprint servers are navigating external pressures and market dynamics, how they are seeking to establish credibility and trust, and how, in doing so, they are reshaping the core functions of scholarly communication.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.024 | 0.093 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.003 | 0.007 |
| Science and technology studies | 0.009 | 0.009 |
| Scholarly communication | 0.033 | 0.019 |
| Open science | 0.002 | 0.010 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.015 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".