Examining gender bias in the feedback shared with family medicine residents
Bibliographic record
Abstract
OBJECTIVES: Competency-based education places increasing emphasis on formative feedback to learners as part of the assessment process. We wished to determine if gender bias was present in the feedback shared with post-graduate medical trainees (residents) in a two-year family medicine residency program at a Canadian university. METHODS: We performed secondary data analyses of documented feedback (FieldNotes) extracted from the Competency-Based Achievement System database. Between 2012 and 2016, 464 preceptors (188 female (F); 276 male (M)) wrote in total 7316 FieldNotes for 192 residents (104 F; 88 M), forming four gender dyads. Descriptive statistics were used to examine trends in FieldNotes frequencies, competencies (Sentinel Habits; SH), progress levels (PL), and the use of adjectives (agentic/competency-based; communal/warmth-based) by preceptors in the FieldNotes. RESULTS: Male and female preceptors wrote on average 7 and 14 FieldNotes, respectively. Female residents received on average more feedback comments from female preceptors (7 notes) than from male preceptors (4 notes). The M-M and M-F resident-preceptor dyads had, respectively, the least and the most 'Stop, Important correction' FieldNotes in both the PGY1 and PGY2 groups. Although preceptors used agentic adjectives more frequently than communal adjectives overall, the F-M resident-preceptor dyad contained the highest proportion of communal adjectives and the lowest proportion of agentic adjectives. CONCLUSIONS: Residents would benefit from multiple opportunities for feedback from both male and female preceptors throughout their residency training. Faculty development to bring attention to potential gender bias may be useful to ensure equitable teaching and quality feedback for learners.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.051 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".