Necessary but Insufficient and Possibly Counterproductive: The Complex Problem of Teaching Evaluations
Bibliographic record
Abstract
The evaluation of clinical teachers' performance has long been a subject of research and debate, yet teaching evaluations (TEs) by students remain problematic. Despite their intuitive appeal, there is little evidence that TEs are associated with students' learning in the classroom or clinical setting. TEs are also subject to many forms of bias and are confounded by construct-irrelevant factors, such as the teacher's physical attractiveness or personality. Yet they are used almost exclusively as evaluations of and feedback to teachers. In this commentary, the authors review the literature on what TEs are meant to do, what they actually do in the real world, and their overall impact. The authors also consider productive ways forward. While TEs are certainly necessary to provide the crucial student voice, they are insufficient as the sole way to assess teachers. Further, they are often counterproductive. TEs carry so much weight for faculty that they can act as a disincentive for teachers to challenge learners and provide them with the critical feedback they often need, lest students give them poor ratings. To address these challenges, changes are needed, including embedding TEs in a programmatic assessment framework. For example, TEs might be used for formative feedback only, while other sources of data, such as peer assessments, learning outcomes, 360-degree feedback, and teacher reflections, could be collated into a portfolio to provide a more meaningful evaluation for teachers. Robust, transparent systems should be in place that dictate how TE data are used and to ensure they are not misused. Clinical teachers who do not "fail to fail" learners but instead take the time and effort to identify and support learners in difficulty should be recognized and rewarded. Learners need this support to succeed and the obligation to protect patients demands it.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.336 | 0.726 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.002 |
| Bibliometrics | 0.007 | 0.005 |
| Science and technology studies | 0.004 | 0.037 |
| Scholarly communication | 0.022 | 0.030 |
| Open science | 0.007 | 0.008 |
| Research integrity | 0.015 | 0.017 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".