[V] Reviewers' Interpretation and Application of Research Quality Criteria in Grant Peer Review
Notice bibliographique
Résumé
Rachel Claus<sup>1</sup> <h4>Objective</h4> Research quality criteria guide grant applications and reviewer evaluations. When research crosses disciplinary boundaries, it requires more expansive quality criteria. Research evaluation is challenging in this context because concepts of research quality are rooted in disciplinary tradition.<sup>1</sup> Reviews are characterized by inconsistency and low agreement<sup>2</sup> and exacerbated by poorly defined criteria.<sup>3</sup> This presentation focuses on how consistently quality criteria are applied and interpreted by reviewers and explores reasons for inconsistency. <h4>Design </h4> All scoring data were collected from 2 noncompetitive review processes of multimillion-dollar research program proposals with aims to improve food security. In 2021, 32 research proposals were evaluated using 17 criteria on a standard 4-point Likert scale (0-3), with limited reviewer overlap. In 2024, 9 proposals were evaluated using 12 criteria, and 4 using 11 criteria, with no reviewer overlap. Each proposal was assessed by a panel of 3 reviewers who scored independently before discussing scores to reach consensus. In total, 45 proposals were evaluated by 80 individual subject matter experts across 45 panels. Score consistency was measured by discrepancy per criterion as the difference between the highest and lowest score awarded in a panel, and how many reviewers agreed on their individual assessments. The standard of consistency was met when at least 2 of 3 reviewers agreed, and the discrepancy between individual reviewer scores was less than or equal to 1 Likert scale point. A total of 696 panel-level measures of consistency were computed. Cross tabulations were used to identify frequencies of inconsistency by criterion. Interviews with reviewers were conducted to understand perceived reasons for individual review discrepancies and disagreement, and individual and panel-level score justifications were analyzed to explore criteria interpretations. <h4>Results</h4> There was a statistically significant relationship between consistency and the evaluation criteria (χ² = 61.013; <i>P</i> < .001). Some criteria were more consistently applied than others. The frequency that each criterion met the standard of consistency is presented in <b>Table 25-1061</b>. Reviewers more frequently applied the following criteria inconsistently when evaluating proposals: comparative advantage (26.7%), monitoring, evaluation, and learning (24.4%), and overall theory of change (24.4%). Justified and transparent costing had the highest rate (71.9%) of inconsistency in 2021 but was not evaluated in the 2024 review cycle. Individual score discrepancies and disagreement were perceived by reviewers to result from diverse disciplinary expertise and a constructive way to achieve comprehensive quality assessments. More problematic reasons for discrepancies included misalignment of criteria to the application and different interpretations resulting from individual values and perceived abilities to make judgments. https://assets.underline.io/markdown_image/1/image/55ce83d15f0173abd5d1a75ab77f6b39.png <h4>Conclusions</h4> The preliminary results indicate scope for criteria clarification to improve consistency in interpretation and reliability of individual reviewer’s quality assessments of proposals. The approach can be adapted to test and inform improvements to quality criteria. <h4>References</h4> 1. Defila R, Di Giulio A. Transdisciplinary development of quality criteria for transdisciplinary research. In: Regeer BJ, Klaassen P, Broerse JEW, eds. <i>Transdisciplinarity for Transformation</i>. Palgrave Macmillan, Cham; 2024. https://doi.org/10.1007/978-3-031-60974-9_5 2. Pier EL, Brauer M, Filut A, et al. Low agreement among reviewers evaluating the same NIH grant applications. <i>Proc Natl Acad Sci U S A</i>. 2018;115(12):2952-2957. doi:10.1073/pnas.1714379115 3. Abdoul H, Perrey C, Amiel P, et al. Peer review of grant applications: criteria used and qualitative study of reviewer practices. <i>PLoS One</i>. 2012;7(9):e46054. doi:10.1371/journal.pone.0046054 <sup>1</sup>Royal Roads University, Victoria, BC, Canada, rachel.claus@royalroads.ca. <h4>Conflict of Interest Disclosures</h4> None reported. <h4>Funding/Support</h4> This research was supported by the Sustainability Research Effectiveness Program, the BC Graduate Scholarship, the Royal Roads Doctoral Scholarship, and the David Harris Flaherty Scholarship. <h4>Role of the Funder/Sponsor</h4> The research was done as part of the Sustainability Research Effectiveness Program, with input provided to the design of the study, review and approval of the abstract and input to the decision to submit the abstract for presentation. Scholarships provided general support without any role in design and conduct of the study; collection, management and interpretation of the data; preparation, review or approval of the abstract, and decision to submit the abstract for presentation.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,099 | 0,032 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,002 | 0,000 |
| Bibliométrie | 0,005 | 0,014 |
| Études des sciences et des technologies | 0,000 | 0,009 |
| Communication savante | 0,000 | 0,001 |
| Science ouverte | 0,002 | 0,001 |
| Intégrité de la recherche | 0,000 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».