Beyond the resume: HR students’ evaluations of interview performances by first and second language speakers
Bibliographic record
Abstract
Purpose High-stakes decision-makers, including human resource (HR) professionals, often exhibit accent biases against second language speakers in professional evaluations. We extend this work by investigating how HR students evaluate simulated job interview performances in English by first and second language speakers of English. Design/methodology/approach Eighty HR students from Calgary and Montreal evaluated the employability of first language (L1) Arabic, English, and Tagalog candidates applying for two positions (nurse, teacher) at four points in the interview (after reading the applicant’s resume, hearing their self-introduction, and listening to each of two responses to interview questions). Candidates’ responses additionally varied in the extent to which they meaningfully answered the interview questions. Findings Students from both cities provided similar evaluations, employability ratings were similar for both advertised positions, and high-quality responses elicited consistently high ratings while evaluations for low-quality responses declined over time. All speakers were evaluated similarly based on their resumes and self-introductions, regardless of their language background. However, evaluations diverged for interview responses, where L1 Arabic and Tagalog speakers were considered more employable than L1 English speakers. Importantly, students’ preference for L1 Arabic and Tagalog candidates over L1 English candidates was magnified when those candidates provided low-quality interview responses. Originality/value Results suggest that even in the absence of dedicated equity, diversity, and inclusion (EDI) training focusing on language and accent bias, HR students may be aware of second language speakers’ potential disadvantages in the workplace, rewarding them in the current evaluations. Findings also highlight the potential influence of contextual factors on HR students’ decision-making.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.003 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.005 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".