MétaCan
Menu
Retour à la cohorte
Enregistrement W7115681829 · doi:10.48448/1gry-m152

Dual Anonymous and Distributed Peer Review for Proposal Review Rankings at the ALMA Observatory

2025· other· W7115681829 sur OpenAlexaboutno aff

Notice bibliographique

RevueUnderline Science Inc. · 2025
Typeother
Langue
Domaine
Thématique
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésComparabilityPrincipal (computer security)Peer reviewProcess (computing)Cohort

Résumé

récupéré en direct d'OpenAlex

John Carpenter,<sup>1</sup> Andrea Corvillón<sup>1</sup> <h4>Objective </h4> In 2021,<sup>1</sup> the Atacama Large Millimeter/submillimeter Array (ALMA) transitioned from single-anonymous panel reviews to dual-anonymous distributed peer review to manage the growing volume of proposal submissions and mitigate potential biases.<sup>2</sup> We conducted a retrospective cohort study to examine associations between this procedural change and proposal rankings across principal investigator (PI) demographic characteristics, using 7 years of data under the previous format (2012-2018) and 4 years under the new process (2021-2024). <h4>Design </h4> We analyzed proposal rankings from 12 ALMA cycles. From 2011 to 2018 (cycles 0-6), proposals were reviewed in topical panels under single-anonymous peer review. In 2019 (cycle 7), investigator lists were randomized while panels were retained. No review process was held in 2020 due to the COVID-19 pandemic. In 2021 (cycle 8), ALMA implemented dual-anonymous, distributed peer review for most proposals. We examined rankings by 3 PI demographic characteristics: (1) experience (number of cycles in which the PI submitted proposals), (2) regional affiliation (Chile, East Asia, Europe, North America, or other), and (3) sex. We grouped proposals by review era: single-anonymous panel review (cycles 1-6; 2012-2018; 9091 proposals) and dual-anonymous distributed peer review (cycles 8-11; 2021-2024; 6490 proposals). Cycle 0 and cycle 7 were excluded from the analysis: cycle 0 because all PIs were, by definition, first-time users of ALMA, and cycle 7 because it was a transitional year for implementing dual anonymity. <h4>Results </h4> Proposal rankings were normalized from 0 (best) to 1 (worst) for comparability across cycles. Proposal counts by demographic subgroup and review era are reported in <b>Table 25-1053</b>. We compared median normalized rankings between review eras using 10,000 bootstrap samples to generate 95% CIs and 2-sided <i>P</i> values. We observed the following associations: (1) for experience, PIs submitting in all cycles had better rankings during single-anonymous panel review than under dual-anonymous distributed review (<i>P</i> = .006), and first-time PIs ranked lowest in both systems with no significant change in rankings (<i>P</i> = .19); (2) for regional affiliation, East Asian PIs showed improved rankings after the transition (<i>P</i> &lt; .001), rankings for European PIs declined (<i>P</i> = .009) but remained above average, and rankings for PIs from North America, Chile, and other regions did not show significant changes (<i>P</i> &gt; .60); and (3) for sex, no statistically significant differences in rankings were observed between male- and female-led proposals in either review system (<i>P</i> = .12). https://assets.underline.io/markdown_image/1/image/ef41fffb9d8404953c23a5d7853259dc.png <h4>Conclusions </h4> Dual-anonymous, distributed peer review was associated with reduced disparities by PI experience and region, consistent with reduced prestige and geographic bias. While increased score variability may have contributed, the nonuniform changes (ie, some groups improved, others remained stable) are inconsistent with a purely noise-driven explanation. These findings suggest that systematic shifts in reviewer behavior, not merely increased randomness, underlie the observed trends. Disentangling the effects of dual anonymity, distributed review, and other concurrent changes (eg, increased use of artificial intelligence) remains an important direction for future research. <h4>References</h4> 1. Donovan Meyer J, Corvillón A, Carpenter JM, et al. Analysis of the ALMA Cycle 8 distributed peer review process. <i>Bull Am Astron Soc</i>. 2022;54(1):43. doi:10.3847/25c2cfeb.4ece85d4 2. Carpenter JM, Corvillón A, Donovan Meyer J, et al. Update on the systematics in the ALMA proposal review process after Cycle 8. <i>Publ Astron Soc Pac</i>. 2022;134:045001. doi:10.1088/1538-3873/ac5b89 <sup>1</sup>Joint ALMA Observatory, Santiago, Chile, john.carpenter@alma.cl. <h4>Conflict of Interest Disclosures</h4> John Carpenter and Andrea Corvillón are employed by the Joint ALMA Observatory, which is jointly managed by Associated Universities Inc/National Radio Astronomy Observatory, the European Organisation for Astronomical Research in the Southern Hemisphere, and the National Astronomical Observatory of Japan on behalf of the ALMA partnership. <h4>Funding/Support</h4> This work was supported by the Joint ALMA Observatory. <h4>Role of the Funder/Sponsor</h4> The funder had a role in the design and conduct of the study; collection, management, analysis, and interpretation of the data; and decision to submit the abstract for presentation. <h4>Acknowledgment</h4> ALMA is a partnership of the European Organization for Astronomical Research in the Southern Hemisphere (representing its member states), National Science Foundation (US), and National Institutes of Natural Sciences (Japan), together with the National Research Council of Canada (Canada), the National Science and Technology Council and Academia Sinica Institute of Astronomy and Astrophysics (Taiwan), and the Korea Astronomy and Space Science Institute (Republic of Korea), in cooperation with the Republic of Chile.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,027
score de la tête « metaresearch » (Gemma)0,018
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche, Méta-épidémiologie (sens strict), Études des sciences et des technologies, Charge utile insuffisante (le modèle a refusé de juger)
Catégories consensuellesMéta-épidémiologie (sens strict), Études des sciences et des technologies, Charge utile insuffisante (le modèle a refusé de juger)
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Synthèse · Signal consensuel: aucune
Score de désaccord entre enseignants0,517
Score d'incertitude au seuil0,999

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0270,018
Méta-épidémiologie (sens strict)0,0020,001
Méta-épidémiologie (sens large)0,0030,001
Bibliométrie0,0010,007
Études des sciences et des technologies0,0030,012
Communication savante0,0010,001
Science ouverte0,0040,003
Intégrité de la recherche0,0010,001
Charge utile insuffisante (le modèle a refusé de juger)0,0050,002

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,040
Tête enseignante GPT0,333
Écart entre enseignants0,293 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.

Devis d'étudeSans objet
Domainenon disponible
GenreSynthèse

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2025
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueUnderline Science Inc.Travaux en français237 207