MétaCan
Menu
Back to cohort
Record W4411445833 · doi:10.7759/cureus.86366

Adjusting for Resident Rater Leniency or Severity Improves the Reliability of Routine Resident Evaluations of Faculty Anesthesiologists

2025· article· en· W4411445833 on OpenAlexaboutno aff
Franklin Dexter, Terrie Vasilopoulos, Brenda G. Fahy

Bibliographic record

VenueCureus · 2025
Typearticle
Languageen
FieldMedicine
TopicRadiology practices and education
Canadian institutionsnot available
Fundersnot available
KeywordsMedicineInter-rater reliabilityReliability (semiconductor)Family medicinePhysical therapyRating scaleStatistics

Abstract

fetched live from OpenAlex

Background The Accreditation Council for Graduate Medical Education (ACGME) of the United States requires all programs to evaluate faculty performance annually. Multiple universities require all faculty to be reviewed annually. These high-stakes evaluations should be reliable. When one anesthesiologist is said to perform better than another, there should be neither frequent Type I errors (i.e., an anesthesiologist is determined to perform better or worse than average when their performance is average) nor Type II errors (i.e., failure to detect above or below average performance). We investigated the generalizability of the finding that if adjustment is not made for rater leniency/severity, results will be statistically unreliable. Methods University of Florida 11-item evaluations were sent on Mondays, over the 2018-19 academic year. 108 ratees (anesthesiologists) had 3302 evaluations by 85 raters (resident physicians). The replicability of the results was assessed by making a comparison with previously published findings from the University of Toronto and the University of Iowa. Results As observed at the University of Toronto, there was greater heterogeneity of scores among raters than among ratees (raters’ eta-squared 0.40; ratees’ 0.22). As observed at the University of Iowa, the Florida rater leniency/severity of scores could not validly be modeled based on a normal distribution, because the distribution of each rater’s mean among raters was not normally distributed (Shapiro-Wilk W = 0.90 (P = 0.00002) among the 75 raters with ≤9 evaluations). Likewise, matching Iowa, Florida’s distribution of each ratee’s mean among ratees was not normally distributed (W = 0.91 (P = 0.00001) among the 94 ratees with ≤9 evaluations). In contrast, treat evaluations with all items scored the maximum as having a value of 1, otherwise 0. As for Iowa, Florida’s corresponding probability distributions of logits were normally distributed (W = 0.99 (P = 0.90) among raters and W = 0.98 (P = 0.09) among ratees, respectively). Rater leniency/severity remained large in the logit scale, with an intraclass correlation coefficient of 0.55. In the original scale, 0/108 ratees had performance that differed significantly from the grand mean of 4.63, using a P < 0.01 criterion. The alternative analysis approach adjusted for the raters’ leniency/severity. Seven ratees were significantly below average (P ≤ 0.0048) and 17 above average (P ≤ 0.0086). Because statistical assumptions were satisfied, analysis in the original scale had a 22% (24/108) false negative rate, like the 21% observed previously at the University of Iowa. Conclusions Routine evaluations of faculty anesthesiologist ratees by anesthesiology resident raters give statistically unreliable results, falsely categorizing performance, unless analyses are adjusted for the covariates of raters. The need for adjustment found with the University of Florida data matches the need for this type of adjustment found at the University of Iowa and the University of Toronto. Thus, this adjustment for raters’ leniency/severity appears to be a general finding for rater/ratee routine evaluations.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.106
metaresearch head score (Gemma)0.286
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Evaluation · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.894
Threshold uncertainty score0.558

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1060.286
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0020.002
Science and technology studies0.0010.001
Scholarly communication0.0020.002
Open science0.0010.002
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.082
GPT teacher head0.437
Teacher spread0.355 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designObservational
DomainEvaluation
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueCureusSame topicRadiology practices and educationFrench-language works237,207