MétaCan
Menu
Retour à la cohorte
Enregistrement W4411445833 · doi:10.7759/cureus.86366

Adjusting for Resident Rater Leniency or Severity Improves the Reliability of Routine Resident Evaluations of Faculty Anesthesiologists

2025· article· en· W4411445833 sur OpenAlexaboutno aff
Franklin Dexter, Terrie Vasilopoulos, Brenda G. Fahy

Notice bibliographique

RevueCureus · 2025
Typearticle
Langueen
DomaineMedicine
ThématiqueRadiology practices and education
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésMedicineInter-rater reliabilityReliability (semiconductor)Family medicinePhysical therapyRating scaleStatistics

Résumé

récupéré en direct d'OpenAlex

Background The Accreditation Council for Graduate Medical Education (ACGME) of the United States requires all programs to evaluate faculty performance annually. Multiple universities require all faculty to be reviewed annually. These high-stakes evaluations should be reliable. When one anesthesiologist is said to perform better than another, there should be neither frequent Type I errors (i.e., an anesthesiologist is determined to perform better or worse than average when their performance is average) nor Type II errors (i.e., failure to detect above or below average performance). We investigated the generalizability of the finding that if adjustment is not made for rater leniency/severity, results will be statistically unreliable. Methods University of Florida 11-item evaluations were sent on Mondays, over the 2018-19 academic year. 108 ratees (anesthesiologists) had 3302 evaluations by 85 raters (resident physicians). The replicability of the results was assessed by making a comparison with previously published findings from the University of Toronto and the University of Iowa. Results As observed at the University of Toronto, there was greater heterogeneity of scores among raters than among ratees (raters’ eta-squared 0.40; ratees’ 0.22). As observed at the University of Iowa, the Florida rater leniency/severity of scores could not validly be modeled based on a normal distribution, because the distribution of each rater’s mean among raters was not normally distributed (Shapiro-Wilk W = 0.90 (P = 0.00002) among the 75 raters with ≤9 evaluations). Likewise, matching Iowa, Florida’s distribution of each ratee’s mean among ratees was not normally distributed (W = 0.91 (P = 0.00001) among the 94 ratees with ≤9 evaluations). In contrast, treat evaluations with all items scored the maximum as having a value of 1, otherwise 0. As for Iowa, Florida’s corresponding probability distributions of logits were normally distributed (W = 0.99 (P = 0.90) among raters and W = 0.98 (P = 0.09) among ratees, respectively). Rater leniency/severity remained large in the logit scale, with an intraclass correlation coefficient of 0.55. In the original scale, 0/108 ratees had performance that differed significantly from the grand mean of 4.63, using a P < 0.01 criterion. The alternative analysis approach adjusted for the raters’ leniency/severity. Seven ratees were significantly below average (P ≤ 0.0048) and 17 above average (P ≤ 0.0086). Because statistical assumptions were satisfied, analysis in the original scale had a 22% (24/108) false negative rate, like the 21% observed previously at the University of Iowa. Conclusions Routine evaluations of faculty anesthesiologist ratees by anesthesiology resident raters give statistically unreliable results, falsely categorizing performance, unless analyses are adjusted for the covariates of raters. The need for adjustment found with the University of Florida data matches the need for this type of adjustment found at the University of Iowa and the University of Toronto. Thus, this adjustment for raters’ leniency/severity appears to be a general finding for rater/ratee routine evaluations.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,001
score de la tête « metaresearch » (Gemma)0,007
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Observationnel · Signal consensuel: Observationnel
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,155
Score d'incertitude au seuil0,825

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0010,007
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,000
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0000,000
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,082
Tête enseignante GPT0,437
Écart entre enseignants0,355 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeObservationnel
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations2
Publié2025
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueCureusMême sujetRadiology practices and educationTravaux en français237 207