Examining the Influence of Gender on Medical Students' Decision Making
Bibliographic record
Abstract
Gender bias, described among practicing physicians, has rarely been examined in medical students. The current study examined the influence of gender bias on medical students' clinical decision making. We experimentally manipulated patient gender in 27 written clinical vignettes embedded in the United States Medical Licensing Examination (USMLE) Step 2 examination (a multiple-choice test of clinical decision making). Female and male patient versions of selected test cases were created within three categories: (1) diseases with previously established evidence of gender bias in the diagnosis or management of the disease, (2) diseases with a higher prevalence in a specific gender, and (3) diseases with similar prevalence in both genders and without evidence of gender bias in the literature. Among the 3059 students who wrote the USMLE Step 2 examination in August 1998, there were small but significant differences in performance on the 12 gender bias cases. Students performed worse for the female patient version of the cases compared with the male patient version of the cases (mean of 55.8% correct for female cases compared with 57.7% correct for male cases) (p < 0. 01). Our data suggest that students were variably influenced by gender bias in their investigation and management of patients in a written test of clinical decision making.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".