MétaCan
Menu
Back to cohort
Record W4312067146 · doi:10.1097/acm.0000000000005006

Necessary but Insufficient and Possibly Counterproductive: The Complex Problem of Teaching Evaluations

2022· article· en· W4312067146 on OpenAlexaff
Shiphra Ginsburg, Lynfa Stroud

Bibliographic record

VenueAcademic Medicine · 2022
Typearticle
Languageen
FieldMedicine
TopicInnovations in Medical Education
Canadian institutionsSunnybrook Health Science CentreThe Wilson CentreSinai Health System
Fundersnot available
KeywordsFormative assessmentConstruct (python library)PsychologyAttractivenessSubject (documents)PortfolioAppealMathematics educationMedical educationSocial psychologyComputer scienceMedicinePolitical science

Abstract

fetched live from OpenAlex

The evaluation of clinical teachers' performance has long been a subject of research and debate, yet teaching evaluations (TEs) by students remain problematic. Despite their intuitive appeal, there is little evidence that TEs are associated with students' learning in the classroom or clinical setting. TEs are also subject to many forms of bias and are confounded by construct-irrelevant factors, such as the teacher's physical attractiveness or personality. Yet they are used almost exclusively as evaluations of and feedback to teachers. In this commentary, the authors review the literature on what TEs are meant to do, what they actually do in the real world, and their overall impact. The authors also consider productive ways forward. While TEs are certainly necessary to provide the crucial student voice, they are insufficient as the sole way to assess teachers. Further, they are often counterproductive. TEs carry so much weight for faculty that they can act as a disincentive for teachers to challenge learners and provide them with the critical feedback they often need, lest students give them poor ratings. To address these challenges, changes are needed, including embedding TEs in a programmatic assessment framework. For example, TEs might be used for formative feedback only, while other sources of data, such as peer assessments, learning outcomes, 360-degree feedback, and teacher reflections, could be collated into a portfolio to provide a more meaningful evaluation for teachers. Robust, transparent systems should be in place that dictate how TE data are used and to ensure they are not misused. Clinical teachers who do not "fail to fail" learners but instead take the time and effort to identify and support learners in difficulty should be recognized and rewarded. Learners need this support to succeed and the obligation to protect patients demands it.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.336
metaresearch head score (Gemma)0.726
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Evaluation · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: Theoretical or conceptual
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.664
Threshold uncertainty score0.818

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.3360.726
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0030.002
Bibliometrics0.0070.005
Science and technology studies0.0040.037
Scholarly communication0.0220.030
Open science0.0070.008
Research integrity0.0150.017
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.044
GPT teacher head0.398
Teacher spread0.354 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designTheoretical or conceptual
DomainEvaluation
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations21
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueAcademic MedicineSame topicInnovations in Medical EducationFrench-language works237,207