Dual Anonymous and Distributed Peer Review for Proposal Review Rankings at the ALMA Observatory
Bibliographic record
Abstract
John Carpenter,<sup>1</sup> Andrea Corvillón<sup>1</sup> <h4>Objective </h4> In 2021,<sup>1</sup> the Atacama Large Millimeter/submillimeter Array (ALMA) transitioned from single-anonymous panel reviews to dual-anonymous distributed peer review to manage the growing volume of proposal submissions and mitigate potential biases.<sup>2</sup> We conducted a retrospective cohort study to examine associations between this procedural change and proposal rankings across principal investigator (PI) demographic characteristics, using 7 years of data under the previous format (2012-2018) and 4 years under the new process (2021-2024). <h4>Design </h4> We analyzed proposal rankings from 12 ALMA cycles. From 2011 to 2018 (cycles 0-6), proposals were reviewed in topical panels under single-anonymous peer review. In 2019 (cycle 7), investigator lists were randomized while panels were retained. No review process was held in 2020 due to the COVID-19 pandemic. In 2021 (cycle 8), ALMA implemented dual-anonymous, distributed peer review for most proposals. We examined rankings by 3 PI demographic characteristics: (1) experience (number of cycles in which the PI submitted proposals), (2) regional affiliation (Chile, East Asia, Europe, North America, or other), and (3) sex. We grouped proposals by review era: single-anonymous panel review (cycles 1-6; 2012-2018; 9091 proposals) and dual-anonymous distributed peer review (cycles 8-11; 2021-2024; 6490 proposals). Cycle 0 and cycle 7 were excluded from the analysis: cycle 0 because all PIs were, by definition, first-time users of ALMA, and cycle 7 because it was a transitional year for implementing dual anonymity. <h4>Results </h4> Proposal rankings were normalized from 0 (best) to 1 (worst) for comparability across cycles. Proposal counts by demographic subgroup and review era are reported in <b>Table 25-1053</b>. We compared median normalized rankings between review eras using 10,000 bootstrap samples to generate 95% CIs and 2-sided <i>P</i> values. We observed the following associations: (1) for experience, PIs submitting in all cycles had better rankings during single-anonymous panel review than under dual-anonymous distributed review (<i>P</i> = .006), and first-time PIs ranked lowest in both systems with no significant change in rankings (<i>P</i> = .19); (2) for regional affiliation, East Asian PIs showed improved rankings after the transition (<i>P</i> < .001), rankings for European PIs declined (<i>P</i> = .009) but remained above average, and rankings for PIs from North America, Chile, and other regions did not show significant changes (<i>P</i> > .60); and (3) for sex, no statistically significant differences in rankings were observed between male- and female-led proposals in either review system (<i>P</i> = .12). https://assets.underline.io/markdown_image/1/image/ef41fffb9d8404953c23a5d7853259dc.png <h4>Conclusions </h4> Dual-anonymous, distributed peer review was associated with reduced disparities by PI experience and region, consistent with reduced prestige and geographic bias. While increased score variability may have contributed, the nonuniform changes (ie, some groups improved, others remained stable) are inconsistent with a purely noise-driven explanation. These findings suggest that systematic shifts in reviewer behavior, not merely increased randomness, underlie the observed trends. Disentangling the effects of dual anonymity, distributed review, and other concurrent changes (eg, increased use of artificial intelligence) remains an important direction for future research. <h4>References</h4> 1. Donovan Meyer J, Corvillón A, Carpenter JM, et al. Analysis of the ALMA Cycle 8 distributed peer review process. <i>Bull Am Astron Soc</i>. 2022;54(1):43. doi:10.3847/25c2cfeb.4ece85d4 2. Carpenter JM, Corvillón A, Donovan Meyer J, et al. Update on the systematics in the ALMA proposal review process after Cycle 8. <i>Publ Astron Soc Pac</i>. 2022;134:045001. doi:10.1088/1538-3873/ac5b89 <sup>1</sup>Joint ALMA Observatory, Santiago, Chile, john.carpenter@alma.cl. <h4>Conflict of Interest Disclosures</h4> John Carpenter and Andrea Corvillón are employed by the Joint ALMA Observatory, which is jointly managed by Associated Universities Inc/National Radio Astronomy Observatory, the European Organisation for Astronomical Research in the Southern Hemisphere, and the National Astronomical Observatory of Japan on behalf of the ALMA partnership. <h4>Funding/Support</h4> This work was supported by the Joint ALMA Observatory. <h4>Role of the Funder/Sponsor</h4> The funder had a role in the design and conduct of the study; collection, management, analysis, and interpretation of the data; and decision to submit the abstract for presentation. <h4>Acknowledgment</h4> ALMA is a partnership of the European Organization for Astronomical Research in the Southern Hemisphere (representing its member states), National Science Foundation (US), and National Institutes of Natural Sciences (Japan), together with the National Research Council of Canada (Canada), the National Science and Technology Council and Academia Sinica Institute of Astronomy and Astrophysics (Taiwan), and the Korea Astronomy and Space Science Institute (Republic of Korea), in cooperation with the Republic of Chile.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.027 | 0.018 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.001 |
| Bibliometrics | 0.001 | 0.007 |
| Science and technology studies | 0.003 | 0.012 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.004 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".