Bibliographic record
Abstract
This year the traditional academic peer review model has come under scrutiny like never before. Does it hinder the rapid dissemination of information vital to clinical and public health decision making? Can it be manipulated? Does it fail to prevent publication of flawed or fraudulent work? It would be difficult to answer any of these questions in the negative, so can we really trust peer review to identify the most relevant and reliable scientific research? This is the question posed for Peer Review Week 2020, marked during Sept 21–25. Preprints, characterised by their publication ahead of peer review, have reached new levels of popularity and prominence this year, and The Lancet Group has recently committed to continuing with its own platform, Preprints with The Lancet, after an initial pilot was received positively.1Kleinert S Horton R on behalf of the Editors of the Lancet GroupPreprints with The Lancet are here to stay.Lancet. 2020; 396: 805Summary Full Text Full Text PDF Scopus (9) Google Scholar But as we wrote in an Editorial earlier this year,2The Lancet Global HealthPublishing in the time of COVID-19.Lancet Glob Health. 2020; 8: e860Summary Full Text Full Text PDF PubMed Scopus (15) Google Scholar despite their advantages in terms of rapid sharing of the potential direction of the answers to time-sensitive research questions, the sheer volume and variability in quality of preprints across a single platform makes it almost impossible for a user to negotiate them meaningfully. Self-serving as it sounds, there is still no proven superior to the tried-and-tested formula of journal-based peer review for assessing the quality of a piece of scientific research—ie, independently chosen external peer reviewers submitting formal critiques, anonymously or not, followed by editor-led decision-making over acceptance or rejection. So what of the problems with slowness, manipulation, and fallibility? The answer, I suggest, lies in the quality of the editorial oversight. As regards timeliness, reviewers could justifiably be forgiven for declining, or deprioritising, invitations to review from a journal that has previously sent very low-quality work. Such work should have been screened out by the editor. Manipulation can be minimised by a keen editorial eye for conflicts of interest and independent verification of reviewers' identity. Flawed and fraudulent work is likely to slip through even the tightest of nets on occasion, particularly when peer review is done at great speed, but the important point here is to have robust editorial mechanisms in place to rectify such issues, resulting in timely correction or retraction. Having a truly diverse and representative reviewer pool is, I suggest, a good way to enhance trust in the peer review process. As part of The Lancet family's ongoing commitments to inclusion and diversity, we analysed our reviewers over the past 12 months by gender and country of origin. Of the 720 individuals who provided at least one review and whose gender we could identify, 288 (40%) were female and 432 (60%) were male. This represents another small improvement in gender balance year-on-year since the 36%/64% split we saw when we first started analysing gender in 2018. In terms of geographic diversity, however, although our reviewers were from 72 different countries overall, almost half were from either the USA or UK, as they were in 2018. The next highest were India on 5%; Australia, China, and Switzerland on 4%; and South Africa and Canada on 3% (appendix pp 1–2). We would expect a degree of skewing, since our choice of reviewers largely reflects research output, and some countries clearly dominate over others in this area. However, the low proportion of Chinese reviewers compared with China's status as the world's largest producer of scientific publications3National Science FoundationPublications output: US trends and international comparisons. National Science Foundation, Alexandria2019https://ncses.nsf.gov/pubs/nsb20206/Date accessed: September 8, 2020Google Scholar is a clear anomaly which we will be exploring over the coming year. Additionally, although we currently make focused efforts to include at least one reviewer from the country or region where a study was done, we believe this is not sufficient. Going forward, we will therefore aim to invite a majority of reviewers from the country or region of interest among our subject specialists. Broadening the geographical diversity of our statistical reviewers is a further area of work for us over the next 12 months. Every Peer Review Week, we publicly name and thank all those who have reviewed for us over the past 12 months (appendix p 3). This year we owe a particularly large debt of gratitude to the reviewers who spared their limited time to review for us, sometimes in a matter of 3 days, during the unprecedented events of early 2020. We know full well that some had been drafted in to perform pandemic-related clinical duties in addition to their own research, had family who became ill, or had small children at home. Reviewers, your dedication to the advancement of science and medicine through helping us to identify the most robust and impactful work has been inspirational and we sincerely thank you. I thank Francesca Cullura for data analysis and figure production. I declare no competing interests. Download .pdf (.7 MB) Help with pdf files Supplementary appendix
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".