Adjusting for Resident Rater Leniency or Severity Improves the Reliability of Routine Resident Evaluations of Faculty Anesthesiologists
Notice bibliographique
Résumé
Background The Accreditation Council for Graduate Medical Education (ACGME) of the United States requires all programs to evaluate faculty performance annually. Multiple universities require all faculty to be reviewed annually. These high-stakes evaluations should be reliable. When one anesthesiologist is said to perform better than another, there should be neither frequent Type I errors (i.e., an anesthesiologist is determined to perform better or worse than average when their performance is average) nor Type II errors (i.e., failure to detect above or below average performance). We investigated the generalizability of the finding that if adjustment is not made for rater leniency/severity, results will be statistically unreliable. Methods University of Florida 11-item evaluations were sent on Mondays, over the 2018-19 academic year. 108 ratees (anesthesiologists) had 3302 evaluations by 85 raters (resident physicians). The replicability of the results was assessed by making a comparison with previously published findings from the University of Toronto and the University of Iowa. Results As observed at the University of Toronto, there was greater heterogeneity of scores among raters than among ratees (raters’ eta-squared 0.40; ratees’ 0.22). As observed at the University of Iowa, the Florida rater leniency/severity of scores could not validly be modeled based on a normal distribution, because the distribution of each rater’s mean among raters was not normally distributed (Shapiro-Wilk W = 0.90 (P = 0.00002) among the 75 raters with ≤9 evaluations). Likewise, matching Iowa, Florida’s distribution of each ratee’s mean among ratees was not normally distributed (W = 0.91 (P = 0.00001) among the 94 ratees with ≤9 evaluations). In contrast, treat evaluations with all items scored the maximum as having a value of 1, otherwise 0. As for Iowa, Florida’s corresponding probability distributions of logits were normally distributed (W = 0.99 (P = 0.90) among raters and W = 0.98 (P = 0.09) among ratees, respectively). Rater leniency/severity remained large in the logit scale, with an intraclass correlation coefficient of 0.55. In the original scale, 0/108 ratees had performance that differed significantly from the grand mean of 4.63, using a P < 0.01 criterion. The alternative analysis approach adjusted for the raters’ leniency/severity. Seven ratees were significantly below average (P ≤ 0.0048) and 17 above average (P ≤ 0.0086). Because statistical assumptions were satisfied, analysis in the original scale had a 22% (24/108) false negative rate, like the 21% observed previously at the University of Iowa. Conclusions Routine evaluations of faculty anesthesiologist ratees by anesthesiology resident raters give statistically unreliable results, falsely categorizing performance, unless analyses are adjusted for the covariates of raters. The need for adjustment found with the University of Florida data matches the need for this type of adjustment found at the University of Iowa and the University of Toronto. Thus, this adjustment for raters’ leniency/severity appears to be a general finding for rater/ratee routine evaluations.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,007 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».