MétaCan
Menu
Back to cohort
Record W3084783588 · doi:10.1016/s2214-109x(20)30416-2

Elements of trust in peer review (and our annual thanks)

2020· article· en· W3084783588 on OpenAlexaboutno aff
Zoë Mullan

Bibliographic record

VenueThe Lancet Global Health · 2020
Typearticle
Languageen
FieldDecision Sciences
TopicAcademic Publishing and Open Access
Canadian institutionsnot available
Fundersnot available
KeywordsScopusScrutinyPopularityPeer reviewNegotiationMedical journalMedicineInternet privacyLibrary sciencePublic relationsPolitical sciencePsychologyMEDLINEComputer scienceLaw

Abstract

fetched live from OpenAlex

This year the traditional academic peer review model has come under scrutiny like never before. Does it hinder the rapid dissemination of information vital to clinical and public health decision making? Can it be manipulated? Does it fail to prevent publication of flawed or fraudulent work? It would be difficult to answer any of these questions in the negative, so can we really trust peer review to identify the most relevant and reliable scientific research? This is the question posed for Peer Review Week 2020, marked during Sept 21–25. Preprints, characterised by their publication ahead of peer review, have reached new levels of popularity and prominence this year, and The Lancet Group has recently committed to continuing with its own platform, Preprints with The Lancet, after an initial pilot was received positively.1Kleinert S Horton R on behalf of the Editors of the Lancet GroupPreprints with The Lancet are here to stay.Lancet. 2020; 396: 805Summary Full Text Full Text PDF Scopus (9) Google Scholar But as we wrote in an Editorial earlier this year,2The Lancet Global HealthPublishing in the time of COVID-19.Lancet Glob Health. 2020; 8: e860Summary Full Text Full Text PDF PubMed Scopus (15) Google Scholar despite their advantages in terms of rapid sharing of the potential direction of the answers to time-sensitive research questions, the sheer volume and variability in quality of preprints across a single platform makes it almost impossible for a user to negotiate them meaningfully. Self-serving as it sounds, there is still no proven superior to the tried-and-tested formula of journal-based peer review for assessing the quality of a piece of scientific research—ie, independently chosen external peer reviewers submitting formal critiques, anonymously or not, followed by editor-led decision-making over acceptance or rejection. So what of the problems with slowness, manipulation, and fallibility? The answer, I suggest, lies in the quality of the editorial oversight. As regards timeliness, reviewers could justifiably be forgiven for declining, or deprioritising, invitations to review from a journal that has previously sent very low-quality work. Such work should have been screened out by the editor. Manipulation can be minimised by a keen editorial eye for conflicts of interest and independent verification of reviewers' identity. Flawed and fraudulent work is likely to slip through even the tightest of nets on occasion, particularly when peer review is done at great speed, but the important point here is to have robust editorial mechanisms in place to rectify such issues, resulting in timely correction or retraction. Having a truly diverse and representative reviewer pool is, I suggest, a good way to enhance trust in the peer review process. As part of The Lancet family's ongoing commitments to inclusion and diversity, we analysed our reviewers over the past 12 months by gender and country of origin. Of the 720 individuals who provided at least one review and whose gender we could identify, 288 (40%) were female and 432 (60%) were male. This represents another small improvement in gender balance year-on-year since the 36%/64% split we saw when we first started analysing gender in 2018. In terms of geographic diversity, however, although our reviewers were from 72 different countries overall, almost half were from either the USA or UK, as they were in 2018. The next highest were India on 5%; Australia, China, and Switzerland on 4%; and South Africa and Canada on 3% (appendix pp 1–2). We would expect a degree of skewing, since our choice of reviewers largely reflects research output, and some countries clearly dominate over others in this area. However, the low proportion of Chinese reviewers compared with China's status as the world's largest producer of scientific publications3National Science FoundationPublications output: US trends and international comparisons. National Science Foundation, Alexandria2019https://ncses.nsf.gov/pubs/nsb20206/Date accessed: September 8, 2020Google Scholar is a clear anomaly which we will be exploring over the coming year. Additionally, although we currently make focused efforts to include at least one reviewer from the country or region where a study was done, we believe this is not sufficient. Going forward, we will therefore aim to invite a majority of reviewers from the country or region of interest among our subject specialists. Broadening the geographical diversity of our statistical reviewers is a further area of work for us over the next 12 months. Every Peer Review Week, we publicly name and thank all those who have reviewed for us over the past 12 months (appendix p 3). This year we owe a particularly large debt of gratitude to the reviewers who spared their limited time to review for us, sometimes in a matter of 3 days, during the unprecedented events of early 2020. We know full well that some had been drafted in to perform pandemic-related clinical duties in addition to their own research, had family who became ill, or had small children at home. Reviewers, your dedication to the advancement of science and medicine through helping us to identify the most robust and impactful work has been inspirational and we sincerely thank you. I thank Francesca Cullura for data analysis and figure production. I declare no competing interests. Download .pdf (.7 MB) Help with pdf files Supplementary appendix

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.010
metaresearch head score (Gemma)0.005
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Commentary · Consensus signal: none
Teacher disagreement score0.825
Threshold uncertainty score0.563

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0100.005
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.001
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0020.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.193
GPT teacher head0.509
Teacher spread0.315 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2020
Admission routes1
Has abstractyes

Explore more

Same venueThe Lancet Global HealthSame topicAcademic Publishing and Open AccessFrench-language works237,207