A peek behind the curtain: Peer review and editorial decision making at <i><scp>S</scp>troke</i>
Bibliographic record
Abstract
Editor's Note The mechanisms of peer review and editorial decision making often appear opaque to junior academic neurologists, especially those who have not yet published many papers or served as journal referees themselves. Previous entries in the NeuroGenesis career development series from the Editor‐in‐Chief have reviewed some of the reasons why faculty should participate as peer reviewers when given the opportunity and the factors that authors should consider in choosing appropriate journals for their own manuscripts. In this article, Sposato et al present the results of a systematic analysis of the editorial process at a leading neurology subspecialty journal; their findings will be of interest to readers at all stages of their careers who seek a better understanding of what goes on “behind the scenes” in journal decisions. — Bernard Chang, MD, NeuroGenesis Editor Objective A better understanding of the manuscript peer‐review process could improve the likelihood that research of the highest quality is funded and published. To this end, we aimed to assess consistency across reviewers' recommendations, agreement between reviewers' recommendations and editors' final decisions, and reviewer‐ and editor‐level factors influencing editorial decisions at the journal Stroke . Methods We analyzed all initial original contributions submitted to Stroke from January 2004 through December 2011. All submissions were linked to the final editorial decision (accept vs reject). We assessed the level of agreement between reviewers (intraclass correlation coefficient). We compared the initial editorial decision (accept, minor revision, major revision, and reject) across reviewers' recommendations. We performed a logistic regression analysis to identify reviewer‐ and editor‐related factors associated with acceptance as the final decision. Results Of 12,902 original submissions to Stroke during the 8‐year study period, the level of agreement between reviewers was between fair and moderate (intraclass correlation coefficient = 0.55, 95% confidence interval [CI] = 0.09–0.75). Likelihood of acceptance was <5% if at least 1 reviewer recommended a rejection. In the multivariate analysis, higher reviewer‐assigned priority scores were related to greater odds of acceptance (odds ratio [OR] = 26.3, 95% CI = 23.2–29.8), whereas higher number of reviewers (OR = 0.54 per additional reviewer, 95% CI = 0.50–0.59) and suggestions for reviewers by authors versus no suggestions (OR = 0.83, 95% CI = 0.73–0.94) had lesser odds of acceptance. Interpretation This analysis of the peer‐review process at Stroke identified several factors that might be targeted to improve the consistency and fairness of the overall process. Ann Neurol 2014;76:151–158
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.034 | 0.226 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.007 | 0.020 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".