MétaCan
Menu
Back to cohort
Record W225686392

New Decision Tool to Evaluate Award Selection Process. (Applied Research)

2002· article· en· W225686392 on OpenAlexaboutno aff
Richard Thornley, Matthew W. Spence, Mark Taylor, Jacques Magnan

Bibliographic record

VenueJournal of Research Administration · 2002
Typearticle
Languageen
FieldHealth Professions
TopicHealth Sciences Research and Education
Canadian institutionsnot available
Fundersnot available
KeywordsGovernment (linguistics)Quality (philosophy)Medical educationPolitical scienceMedicine
DOInot available

Abstract

fetched live from OpenAlex

Introduction Established by the Government of Alberta in 1979, the Alberta Heritage Foundation for Medical Research (AHFMR) supports health research at Alberta universities and other research-related institutions. The foundation supports nearly 230 faculty-level researchers recruited from Alberta and around the world, and approximately 500 researchers-in-training (i.e., summer students, graduate students, and post-doctoral fellows, collectively known as trainees). The AHFMR's gross expenditure for fiscal year (FY) 2000-2001 was approximately $53 million, of which $6.7 million (12.6%) was committed to the funding of trainees. (1) This article describes the foundation's initiative to improve the peer review process for its competitive training awards. Peer review is frequently used for both ex ante and ex post evaluation of the quality of the scientific enterprise (Geisler, 2000; Kostoff, 1992; Luukkonen-Gronow 1987; United States General Accounting Office, 1997). Ex ante evaluation assesses quality in advance of performance, as in the case of applications for research funding. Conversely, ex post evaluation assesses quality retrospectively, as in the case of papers submitted to scientific journals. The case described here entails ex ante review of applications for funding, to anticipate the future performance of research trainees. The AHFMR's original review process for training award applications considered three general criteria: (a) the quality of the candidate, (b) the appropriateness of the proposed research environment, and (c) the merit of the proposed research project. Applications were rated following a multiple-step committee process on a scale of 0 to 5, the single score representing an aggregation of performance in relation to all criteria. Zero is considered an unacceptable application whereas a score of 5 is an outstanding application. This approach was used by the foundation to review applications for its training awards until the end of FY2000, when the foundation piloted the new process described here. Geisler (2000) suggested that peer review should be well-defined, rational, fair, timely, cost-effective, anonymous, and responsive. While most of these general characteristics were reflected in the AHFMR's original review process for its training awards, a number of specific issues provided the incentive for the foundation to try to improve the process. First, the number of proposals submitted was increasing and there was a need to more efficiently evaluate them. In FY1997, the AHFMR received 182 applications for full-time studentships, as compared to 276 in FY2000 and 307 in FY2001. This resulted in the need for more reviewers, most of whom were reporting that they had increasingly less time to devote to such activities. Also, the increase in proposals meant that committees were faced with extending the duration of their meetings or spending less time reviewing each application, neither of which was considered to be a desirable alternative. This issue was complicated by an increase in turnover on the foundation's review committees. In general, this may have been in response to reviewer fatigue, a recent and widespread phenomenon in the research funding sector resulting from a proliferation of requests to individuals to sit on review panels (Brzustowski, 2000a; Brzustowski, 2000b; Cunningham, Boden, Glynn, & Hills, 2001; Smith, 2001). There was a sense that turnover resulted in less consistency in the application of criteria within and between competitions, and an increased administrative burden in recruiting and training committee members. Two trends relating to scores awarded to applications also influenced the AHFMR's decision to redesign its review process. In theory, the overall score awarded to each application represented an integration of all parts of the application; however, in practice each reviewer's interpretation resulted in variable weighting of different criteria. …

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.092
metaresearch head score (Gemma)0.261
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Incentives · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.908
Threshold uncertainty score0.486

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0920.261
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0030.003
Bibliometrics0.0150.012
Science and technology studies0.0020.002
Scholarly communication0.0100.007
Open science0.0030.004
Research integrity0.0030.004
Insufficient payload (model declined to judge)0.0370.007

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.446
GPT teacher head0.635
Teacher spread0.189 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designTheoretical or conceptual
DomainIncentives
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2002
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Research AdministrationSame topicHealth Sciences Research and EducationFrench-language works237,207