Gender Representation of Health Care Professionals in Large Language Model–Generated Stories
Bibliographic record
Abstract
Importance: With the growing use of large language models (LLMs) in education and health care settings, it is important to ensure that the information they generate is diverse and equitable, to avoid reinforcing or creating stereotypes that may influence the aspirations of upcoming generations. Objective: To evaluate the gender representation of LLM-generated stories involving medical doctors, surgeons, and nurses and to investigate the association of varying personality and professional seniority descriptors with the gender proportions for these professions. Design, Setting, and Participants: This is a cross-sectional simulation study of publicly accessible LLMs, accessed from December 2023 to January 2024. GPT-3.5-turbo and GPT-4 (OpenAI), Gemini-pro (Google), and Llama-2-70B-chat (Meta) were prompted to generate 500 stories featuring medical doctors, surgeons, and nurses for a total 6000 stories. A further 43 200 prompts were submitted to the LLMs containing varying descriptors of personality (agreeableness, neuroticism, extraversion, conscientiousness, and openness) and professional seniority. Main Outcomes and Measures: The primary outcome was the gender proportion (she/her vs he/him) within stories generated by LLMs about medical doctors, surgeons, and nurses, through analyzing the pronouns contained within the stories using χ2 analyses. The pronoun proportions for each health care profession were compared with US Census data by descriptive statistics and χ2 tests. Results: In the initial 6000 prompts submitted to the LLMs, 98% of nurses were referred to by she/her pronouns. The representation of she/her for medical doctors ranged from 50% to 84%, and that for surgeons ranged from 36% to 80%. In the 43 200 additional prompts containing personality and seniority descriptors, stories of medical doctors and surgeons with higher agreeableness, openness, and conscientiousness, as well as lower neuroticism, resulted in higher she/her (reduced he/him) representation. For several LLMs, stories focusing on senior medical doctors and surgeons were less likely to be she/her than stories focusing on junior medical doctors and surgeons. Conclusions and Relevance: This cross-sectional study highlights the need for LLM developers to update their tools for equitable and diverse gender representation in essential health care roles, including medical doctors, surgeons, and nurses. As LLMs become increasingly adopted throughout health care and education, continuous monitoring of these tools is needed to ensure that they reflect a diverse workforce, capable of serving society's needs effectively.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".