Quality In-Training Evaluation Reports—Does Feedback Drive Faculty Performance?
Bibliographic record
Abstract
PURPOSE: Clinical faculty often complete in-training evaluation reports (ITERs) poorly. Faculty development (FD) strategies should address this problem. An FD workshop was shown to improve ITER quality, but few physicians attend traditional FD workshops. To reach more faculty, the authors developed an "at-home" FD program offering participants various types of feedback on their ITER quality based on the workshop content. Program impact is evaluated here. METHOD: Ninety-eight participants from four medical schools, all clinical supervisors, were recruited in 2009-2010; 37 participants completed the study. These were randomized into five groups: a control group and four other groups with different feedback conditions. ITER quality was assessed by two raters using a validated tool: the completed clinical evaluation report rating (CCERR). Participants were given feedback on their ITER quality based on group assignment. Six months later, participants submitted new ITERs. These ITERs were assessed using the CCERR, and feedback was sent to participants on the basis of their group assignment. This process was repeated two more times, ending in 2012. RESULTS: CCERR scores from the participants in all feedback groups were collapsed (n=27) and compared with scores from the control group (n=10). Mean CCERR scores significantly increased over time for the feedback group but not the control group. CONCLUSIONS: The results suggest that faculty are able to improve ITER quality following a minimal "at-home" FD intervention. This also adds to the growing literature that has found success with improving the quality of trainee assessments following rater training.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.009 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".