MétaCan
Menu
Back to cohort
Record W4396882454 · doi:10.48550/arxiv.2405.06563

What Can Natural Language Processing Do for Peer Review?

2024· preprint· en· W4396882454 on OpenAlexafffund
Ilia Kuznetsov, Osama Mohammed Afzal, Koen Dercksen, Nils Dycke, Alexander Goldberg, Tom Hope, Dirk Hovy, Jonathan Kummerfeld, Anne Lauscher, Kevin Leyton‐Brown, Sheng Lu, Prof. Mausam, Margot Mieskes, Aurélie Névéol, Danish Pruthi, Lizhen Qu, Roy Schwartz, Noah A. Smith, Thamar Solorio, Jingyan Wang, Xiaodan Zhu, Anna Rogers, Nihar B. Shah, Iryna Gurevych

Bibliographic record

VenueTUbilio (Technical University of Darmstadt) · 2024
Typepreprint
Languageen
FieldComputer Science
TopicTopic Modeling
Canadian institutionsQueen's UniversityOkanagan University CollegeUniversity of British Columbia, Okanagan CampusUniversity of British Columbia
FundersOffice of Naval ResearchBundesministerium für Bildung und ForschungDeutsche ForschungsgemeinschaftNatural Sciences and Engineering Research Council of CanadaEuropean CommissionAustralian GovernmentAlberta Machine Intelligence InstituteCanadian Institute for Advanced ResearchNational Science Foundation
KeywordsOperationalizationComputer sciencePaceProcess (computing)Peer reviewArtificial intelligenceTechnical peer reviewField (mathematics)AsideData sciencePolitical scienceLinguistics

Abstract

fetched live from OpenAlex

The number of scientific articles produced every year is growing rapidly. Providing quality control over them is crucial for scientists and, ultimately, for the public good. In modern science, this process is largely delegated to peer review -- a distributed procedure in which each submission is evaluated by several independent experts in the field. Peer review is widely used, yet it is hard, time-consuming, and prone to error. Since the artifacts involved in peer review -- manuscripts, reviews, discussions -- are largely text-based, Natural Language Processing has great potential to improve reviewing. As the emergence of large language models (LLMs) has enabled NLP assistance for many new tasks, the discussion on machine-assisted peer review is picking up the pace. Yet, where exactly is help needed, where can NLP help, and where should it stand aside? The goal of our paper is to provide a foundation for the future efforts in NLP for peer-reviewing assistance. We discuss peer review as a general process, exemplified by reviewing at AI conferences. We detail each step of the process from manuscript submission to camera-ready revision, and discuss the associated challenges and opportunities for NLP assistance, illustrated by existing work. We then turn to the big challenges in NLP for peer review as a whole, including data acquisition and licensing, operationalization and experimentation, and ethical issues. To help consolidate community efforts, we create a companion repository that aggregates key datasets pertaining to peer review. Finally, we issue a detailed call for action for the scientific community, NLP and AI researchers, policymakers, and funding bodies to help bring the research in NLP for peer review forward. We hope that our work will help set the agenda for research in machine-assisted scientific quality control in the age of AI, within the NLP community and beyond.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.427
metaresearch head score (Gemma)0.707
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Evaluation · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.573
Threshold uncertainty score0.707

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.4270.707
Meta-epidemiology (narrow)0.0030.002
Meta-epidemiology (broad)0.0070.004
Bibliometrics0.0190.016
Science and technology studies0.0190.026
Scholarly communication0.0570.098
Open science0.0130.018
Research integrity0.0230.020
Insufficient payload (model declined to judge)0.0270.044

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.023
GPT teacher head0.282
Teacher spread0.259 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designTheoretical or conceptual
DomainEvaluation
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations5
Published2024
Admission routes2
Has abstractyes

Explore more

Same venueTUbilio (Technical University of Darmstadt)Same topicTopic ModelingFrench-language works237,207