Ethnic and Gender Biases in Clinical Performance Assessment (CPA) in Healthcare Education: A Systematic Review
Bibliographic record
Abstract
Abstract Background: Several forms of bias, including ethnic and gender bias, are thought to impact evaluations on Clinical Performance Assessments (CPAs). Unfairness may influence student learning attitudes if a loss of trust causes a lack of engagement in learning. Understanding the biases occurring in CPAs can lead to well-designed examiner training to ensure equality and fairness. The purpose of this systematic review is to determine the current evidence in the literature for ethnic and/or gender bias by examiners evaluating pre-licensure healthcare students in CPAs using standardized patients (SPs). Methods: Literature was systematically searched in CINAHL, PubMed and Medline from inception to February 2019, and no date range was set. Studies related to the investigation of ethnic and/or gender biases occurring in CPAs using SPs for examining health professions students were selected. A systematic review was conducted to assess the methodological quality and strength of evidence of relevant research and to identify if any potential ethnic and/or gender bias occurred in CPAs. The Guidelines for Critical Review were used to appraise the selected studies. Results: Nine studies published from 2003 to 2017 were retrieved for review. Three studies met all the Guidelines for Critical Review quality criteria, indicating stronger evidence of their outcomes, two of the studies reported ethnic and/or gender bias existing in the CPAs. Overall, four studies found ethnic and/or gender bias in CPAs, but all study results had small effect sizes. Conclusions: No systematic and consistent bias was found across the studies; nonetheless, the possibility of ethnic or gender bias by some examiners cannot be ignored. To minimize potential examiner bias, the investigation of Frame of Reference training, multiple examiners per station, and combination assessments in CPAs is recommended.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.042 | 0.229 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.007 | 0.008 |
| Bibliometrics | 0.010 | 0.012 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.005 | 0.005 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".