MétaCan
Menu
Back to cohort
Record W7115739120 · doi:10.48448/4k9r-hb48

[V] Reviewers' Interpretation and Application of Research Quality Criteria in Grant Peer Review

2025· other· W7115739120 on OpenAlexaffabout

Bibliographic record

VenueUnderline Science Inc. · 2025
Typeother
Language
Field
Topic
Canadian institutionsRoyal Roads University
Fundersnot available
KeywordsLikert scaleContext (archaeology)Consistency (knowledge bases)Presentation (obstetrics)Quality (philosophy)Scale (ratio)Research design

Abstract

fetched live from OpenAlex

Rachel Claus<sup>1</sup> <h4>Objective</h4> Research quality criteria guide grant applications and reviewer evaluations. When research crosses disciplinary boundaries, it requires more expansive quality criteria. Research evaluation is challenging in this context because concepts of research quality are rooted in disciplinary tradition.<sup>1</sup> Reviews are characterized by inconsistency and low agreement<sup>2</sup> and exacerbated by poorly defined criteria.<sup>3</sup> This presentation focuses on how consistently quality criteria are applied and interpreted by reviewers and explores reasons for inconsistency. <h4>Design </h4> All scoring data were collected from 2 noncompetitive review processes of multimillion-dollar research program proposals with aims to improve food security. In 2021, 32 research proposals were evaluated using 17 criteria on a standard 4-point Likert scale (0-3), with limited reviewer overlap. In 2024, 9 proposals were evaluated using 12 criteria, and 4 using 11 criteria, with no reviewer overlap. Each proposal was assessed by a panel of 3 reviewers who scored independently before discussing scores to reach consensus. In total, 45 proposals were evaluated by 80 individual subject matter experts across 45 panels. Score consistency was measured by discrepancy per criterion as the difference between the highest and lowest score awarded in a panel, and how many reviewers agreed on their individual assessments. The standard of consistency was met when at least 2 of 3 reviewers agreed, and the discrepancy between individual reviewer scores was less than or equal to 1 Likert scale point. A total of 696 panel-level measures of consistency were computed. Cross tabulations were used to identify frequencies of inconsistency by criterion. Interviews with reviewers were conducted to understand perceived reasons for individual review discrepancies and disagreement, and individual and panel-level score justifications were analyzed to explore criteria interpretations. <h4>Results</h4> There was a statistically significant relationship between consistency and the evaluation criteria (χ² = 61.013; <i>P</i> &lt; .001). Some criteria were more consistently applied than others. The frequency that each criterion met the standard of consistency is presented in <b>Table 25-1061</b>. Reviewers more frequently applied the following criteria inconsistently when evaluating proposals: comparative advantage (26.7%), monitoring, evaluation, and learning (24.4%), and overall theory of change (24.4%). Justified and transparent costing had the highest rate (71.9%) of inconsistency in 2021 but was not evaluated in the 2024 review cycle. Individual score discrepancies and disagreement were perceived by reviewers to result from diverse disciplinary expertise and a constructive way to achieve comprehensive quality assessments. More problematic reasons for discrepancies included misalignment of criteria to the application and different interpretations resulting from individual values and perceived abilities to make judgments. https://assets.underline.io/markdown_image/1/image/55ce83d15f0173abd5d1a75ab77f6b39.png <h4>Conclusions</h4> The preliminary results indicate scope for criteria clarification to improve consistency in interpretation and reliability of individual reviewer’s quality assessments of proposals. The approach can be adapted to test and inform improvements to quality criteria. <h4>References</h4> 1. Defila R, Di Giulio A. Transdisciplinary development of quality criteria for transdisciplinary research. In: Regeer BJ, Klaassen P, Broerse JEW, eds. <i>Transdisciplinarity for Transformation</i>. Palgrave Macmillan, Cham; 2024. https://doi.org/10.1007/978-3-031-60974-9_5 2. Pier EL, Brauer M, Filut A, et al. Low agreement among reviewers evaluating the same NIH grant applications. <i>Proc Natl Acad Sci U S A</i>. 2018;115(12):2952-2957. doi:10.1073/pnas.1714379115 3. Abdoul H, Perrey C, Amiel P, et al. Peer review of grant applications: criteria used and qualitative study of reviewer practices. <i>PLoS One</i>. 2012;7(9):e46054. doi:10.1371/journal.pone.0046054 <sup>1</sup>Royal Roads University, Victoria, BC, Canada, rachel.claus@royalroads.ca. <h4>Conflict of Interest Disclosures</h4> None reported. <h4>Funding/Support</h4> This research was supported by the Sustainability Research Effectiveness Program, the BC Graduate Scholarship, the Royal Roads Doctoral Scholarship, and the David Harris Flaherty Scholarship. <h4>Role of the Funder/Sponsor</h4> The research was done as part of the Sustainability Research Effectiveness Program, with input provided to the design of the study, review and approval of the abstract and input to the decision to submit the abstract for presentation. Scholarships provided general support without any role in design and conduct of the study; collection, management and interpretation of the data; preparation, review or approval of the abstract, and decision to submit the abstract for presentation.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.099
metaresearch head score (Gemma)0.032
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Meta-epidemiology (narrow), Science and technology studies, Insufficient payload (model declined to judge)
Consensus categoriesMetaresearch, Insufficient payload (model declined to judge)
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Other design · Consensus signal: none
GenreCandidate signal: Review · Consensus signal: none
Teacher disagreement score0.687
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0990.032
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0020.000
Bibliometrics0.0050.014
Science and technology studies0.0000.009
Scholarly communication0.0000.001
Open science0.0020.001
Research integrity0.0000.002
Insufficient payload (model declined to judge)0.0010.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.109
GPT teacher head0.510
Teacher spread0.401 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designOther design
Domainnot available
GenreReview

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes2
Has abstractyes

Explore more

Same venueUnderline Science Inc.French-language works237,207