Asynchronous Video Interviews in Recruitment and Selection: Lights, Camera, Action!
Bibliographic record
Abstract
Technological developments have rapidly changed how job applicants and organizations interact (Woods et al. 2019). One of the most influential recent advances in assessment methods is the Asynchronous Video Interview (AVI) (Lukacik et al. 2022). When completing an AVI, applicants connect to an online platform and use a device with a camera and microphone (e.g., computer or smartphone) to video-record responses to interview questions. The video-recorded responses are either manually reviewed by human judges or automatically via artificial intelligence (AI) algorithms. AVIs afford recruiters considerable convenience, scale, and reach benefits, while applicants benefit from the flexibility of completing AVIs at a time and place of their choosing (Basch and Melchers 2019; Guchait et al. 2014). With these benefits, it is unsurprising that recruiters around the world have adopted AVIs, with social distancing requirements during the COVID-19 pandemic accelerating AVI adoption (Dunlop et al. 2022). As practitioners develop and adopt new assessment methods, there becomes an urgent need for an evidence base informing these methods' efficacy. The International Journal of Selection and Assessment responded to this need with a special issue focused on research relating to all aspects of AVIs. The call for submissions on AVI research garnered one of the largest responses to a special issue in the journal's history and resulted in 15 articles accepted for publication, highlighting the interest in AVIs among the assessment and selection research community. In this editorial, we identify and summarize five themes emerging from these papers, namely AVI content (design features and questions), communication around AVIs, impression management behavior in AVIs, fairness in AVI evaluations, and automated scoring. We conclude by sharing our reflections. Although it is convenient to use the term “AVI” to describe a method holistically, much like in-person job interviews, all AVIs are not the same. Rather, the seminal review by Lukacik et al. (2022) highlighted the many AVI design features that can be adjusted by recruiters or AVI vendors and explained how the different settings might affect important outcomes such as applicant reactions (e.g., fairness perceptions, anxiety) and behaviors (e.g., performance and impression management). Similarly, AVIs contain widely differing interview questions, and the question content and way in which the questions are asked may affect the information that candidates divulge. The fact that more than half of the articles forming this special issue involved the study of AVI content reflects the strong interest in the impact that AVI design choices have on outcomes. Lukacik and Bourdage (2025) shared the results of three studies, each of which experimentally varied two AVI features, including: response preparation time and visibility of the interviewee's “self-view” video feed (Study 1), the opportunity to review and re-record a response (Study 2), and the presence of warnings not to be dishonest and information about whether the responses would be human- versus AI-evaluated (Study 3). Although these authors recruited large samples and investigated a range of behavioral outcomes and applicant reactions, in general, they found very little evidence that design features affect candidate reactions and behaviors. Longer response preparation time resulted in higher interview ratings, and some evidence suggested that affording re-recording opportunities improved performance. However, these authors discovered that interviewees who submitted a re-recorded response tended to be evaluated less favorably than participants who did not re-record—perhaps suggesting that lower performing interviewees are most likely to feel the need to re-record. Four articles examined the effects of incorporating richer media (than text alone) into an AVI, for example, during a “welcome” message or when asking the interview questions. Two of these studies (Niemitz et al. 2024; Salimian Rizi and Roulin 2023) directly compared rich video to text content. In a lab experiment, Salimian Rizi and Roulin (2023) discovered that video introductions and questions, compared with text only, increased perceptions of social presence, which triggered greater honest impression management, lower levels of interview anxiety, and stronger interview performance. By contrast, in a field experiment with real applicants, Niemitz et al. (2024) did not find similar improvements on perceptions of social presence, performance, impression management, or applicant reactions for video versus text questions. Although Niemitz et al.'s (2024) study involved a smaller sample and also included a video-based introduction in both conditions, the mean differences they observed were generally opposite to expectations, suggesting that research into the value of video media when questioning requires further attention. In the third article, Moore et al. (2024) also observed contradictory results across two studies of rich media. These authors adopted a self-determination theory (Deci and Ryan 2000) lens to understand whether the content of video messaging can support interviewees' need for relatedness. Moore et al. (2024) first showed that including an empathic and humorous video introduction, versus a more “professional” introduction, was associated with more positive reactions—including relatedness need-satisfaction. However, this sample did not complete an AVI after watching the introductory video. In their second sample of participants who did complete an AVI after watching the video introductions, the positive reactions were not replicated. Put together, it seems that respondents generally do appreciate humorous messaging, however, that appreciation is short-lived, with the experience of completing the interview ‘washing away’ those positive effects. Rich media can take on forms other than video, and in the fourth study of richer media, Min et al. (2024) examined reactions to AI-generated ‘avatars’ who might play the role of a virtual interviewer in an AVI. Using a vignette design, Min et al. found that an AI avatar's appearance and communication style, along with the depth of feedback provided by the avatar, can impact interpersonal and informational justice perceptions. This initial investigation implies that virtual interviewers are a possible alternative to text- or video-based content delivery, however, the design of the virtual interviewer may not be a simple superficial decision. The special issue includes two papers that focused on the content of AVI questions. Patel et al. (2025) investigated how adding two targeted follow-up questions, or probes, to each question affected applicant reactions and behavior. Such probes allow interviewees to provide additional information that they may have overlooked in their initial response. Across two studies, Patel et al. found that probes, when offered, were used by a very high proportion of interviewees, and the presence of probes caused stronger interview performance. Thus, much like in face-to-face interviews, relevant probes may increase the amount of criterion-relevant information collected in AVIs. In the second paper, Benson et al. (2025) addressed the challenge of improving the equitability of AVIs for autistic candidates. Partnering with an autism employment advising organization to recruit autistic job seekers, these researchers analyzed how subtle interview question wording changes (e.g., reducing metaphors) affected AI interview scores. However, the team found that the question modifications failed to reduce AI score differences between autistic and neurotypical job seekers. Further, a textual analysis revealed that autistic candidates spoke fewer words overall and used fewer expressions about influence, power, and confidence, and these differences partly explained the gap in AI scores. The authors conclude that well-intentioned adjustments of AVI questions do not fully resolve the interview-related disadvantages faced by autistic job seekers, implying that other approaches may be necessary to ensure equitable access to jobs for autistic candidates. Like all selection instruments, AVIs are usually embedded in a broader selection system. Many AVI platforms allow organizations to communicate with candidates through the AVI platform directly or via Applicant Tracking System (ATS) integration. Three studies showed that the content of these communications can affect applicant reactions or behaviors. Falls et al. (2025) investigated whether the explanations provided in rejection letters, following an AVI with automated evaluation, could improve applicants' reactions to the rejection. Specifically, they provided rejection letters to participants that explained the benefits of consistency in automatically scored AVIs (e.g., all interviews are scored the same way, so are not affected by human raters getting tired), or that pointed out the benefits in terms of being able to have control over the process (e.g., when/where the interview takes place), a combination of both explanations, or no explanation. They found that only a combined explanation had a significant but small effect on interviewees' perceptions of procedural fairness. In their communications with applicants, employers might consider providing information to candidates about the person who will evaluate their responses, thus not leaving applicants wondering whether anybody is watching. Orji et al. (2025) discovered that if an evaluator's LinkedIn profile is shared before an AVI, interviewees tended to engage in more other-focused impression management, particularly ingratiation, compared to interviewees presented with no information about the evaluator. The rationale here is that the LinkedIn profile provides interviewees with a ‘target’ to try to influence through signaling person-organization fit. In a novel exploratory analysis of the interviewees' stated goals, however, Orji et al. found that the increased focus on ingratiation, triggered by the exposure to evaluator information, may have cost some interviewees opportunities to communicate person-job fit information. While organizations can make many design and communication decisions that impact AVI perceptions and behaviors, many circumstances exist outside organizational control that can also meaningfully impact these outcomes. Two studies examined how circumstances outside of organizational control affect impression management in AVIs. One of these studies was a multi-national investigation into the impact of culture on self-reported impression management behavior and AVI performance (Arseneault and Roulin 2024). Such investigations are extremely timely and important when a major affordance of AVIs is the reach of employers to candidates from across the globe. These authors observed many nuanced relationships with cultural dimensions and impression management behavior, some of which ran counter to their hypotheses. For example, interviewees from cultures higher on assertiveness appeared to be less likely to engage in self-focused honest impression management and other-focused deceptive impression management. Further, a culture's gender egalitarianism and uncertainty avoidance appeared not to influence any form of impression management. Overall, while observing some differences between cultures, Arseneault and Roulin (2024) found high levels of honest and self-promotion impression management across cultures in AVIs. In the second study of impression management, Canagasuriam and Lukacik (2024) investigated the implications of interviewees' use of large language models (LLMs), a type of generative AI. In an example of what Lievens and Dunlop (2025) characterized as using generative AI as a “substitute” [for the interviewee], the interviewees in Canagasuriam and Lukacik's (2024) two experimental conditions were instructed to enter the questions into an LLM and read the LLM-generated response verbatim or to personalize the LLM-generated response when responding. The study revealed that responses from LLM users were evaluated substantially more positively than responses from non-users, suggesting that using LLMs as a substitute to complete an assessment could garner large advantages to candidates (see also Hickman et al. 2024). Canagasuriam and Lukacik also found, however, that evaluators regarded the responses from the non-users as more honest, and a content analysis of the LLM-generated responses revealed a high degree of consistency in content across interviewees. Thus, while there is clearly some urgency to better understand whether and when AVI validity is threatened by the availability of generative AI to candidates, it may be that over time, the simple repetition of LLM-generated content by candidates will become more easily detected by employers. We highlighted in our call for papers that AVIs can introduce new fairness concerns. In particular, video recordings from AVIs can contain many sources of information in the background, most of which are likely irrelevant to job performance but could nonetheless influence evaluations (Torres and Gregory 2018). Two studies published in this special issue focused on how background information affected evaluations of interviewees' AVI responses. In one study, Springle and Bourdage (2025) examined how candidate evaluations are affected by lower versus high socioeconomic status (SES) signals in the background and evaluator cognitive load. Although the authors observed an association between the perceptions of a candidates' SES and hireability, the SES cues did not directly affect hireability, regardless of the evaluators' cognitive load. In the other study, Basch et al. (2024) found in a German sample some evidence that the presence of religious (specifically, Islamic) paraphernalia in the background had a small negative effect on evaluations of competence (but not on overall ratings or warmth), whereas background signals of same-sex romantic attraction did not affect evaluations. Finally, key to the scalability of AVIs is the advent of automated scoring algorithms, and two studies in this special issue investigated the efficacy of these approaches. Juničić et al. (2025) examined the efficacy of a multi-modal machine learning algorithm in detecting deception during an AVI. The algorithm drew from paraverbal cues, verbal cues, nonverbal cues, and facial expressions, collected on camera as participants responded to interview questions. Although this study involved a physical interviewer, these types of cues and information are also present in AVIs (Hickman et al. 2022; Koutsoumpis et al. 2024), as are automated scoring methods that can draw from them (Liff et al. 2024). It is important, therefore, to ascertain whether such information is diagnostic of candidate deception. Using a within-subjects design, where interviewees completed the interview twice under honest and ideal candidate instructions, Juničić et al. (2025) found that participants reported being less honest when responding as a job candidate. The authors' algorithms for detecting which experimental condition interviewees were in performed better than chance—but accuracy was not sufficient to justify using deception detection algorithms in practice. Stevenor et al. (2024) examined the validity of AI-based personality scores with real job applicants. The authors trained machine learning models on “low-stakes” interviews (i.e., mock interview samples) to predict the Big Five traits, then applied these models to “high-stakes” applicant interviews. The AI scores converged moderately with human interviewer ratings. However, the authors found weak discriminant validity (i.e., overlap among the trait scores) and stronger correlations with verbal behaviors—particularly word count—than expected. These findings highlight that algorithms may capture key personality indicators observed by human raters and can be transposed from low-stakes to high-stakes environments but also underscore the persistent challenges of ensuring adequate construct differentiation. Lastly, Stevenor et al. (2024) reported criterion validity evidence, but on a very small sample. Consequently, knowledge of AVIs' criterion validity for real job applicants remains elusive. We conclude by contemplating the implications for the volume of work presented in this special issue and the future of research into AVI, identifying areas we see as being in most need of further attention. Though there were some exceptions, we were surprised by the general pattern suggesting that changing the design features of AVIs had relatively small or nonreplicable effects on applicant reactions and behaviors. A simple explanation is that the features are indeed trivial relative to the effects of an AVI experience itself. There may, however, be alternatives that warrant consideration. The use of between-subjects experimental designs means that participants do not experience a counterfactual and thus may lack a clear reference point for judging the AVI experience they were assigned to. For example, would an interviewee react positively to an opportunity to re-record a response if it never occurs to them that this option could be switched off? Over time, however, we expect job seekers will be exposed to a greater number and variety of AVIs and may eventually better recognize the AVI features that support or frustrate them. Thus, we may effects in between-subjects studies of AVI features in the future research could our by using designs to how applicants would feel when they are of this the generally small effects of features on such as performance, impression management, or anxiety, that less It is also possible that the effects of some AVI features are but only such that they are in the research could capture the in the benefits of AVI features directly by applicant reactions or indicators such as during the but reactions might not be A third is that some features have very effects but only for a interviewees, and also only for a small of the AVI. For example, an interviewee who was while a response to an AVI question would likely appreciate the opportunity to re-record and that interviewee's to may be if re-recording was in that but interviewees are not to on the availability of the re-record Finally, as by Falls et al. who showed that explanations about AI can improve reactions to rejection we that it may be for AVI vendors and users to to applicants AVI features are being and the benefits of those of the questions AVIs to be relatively and so we were by the two papers that focused on questions and question design to improve for autistic job candidates. The of these two investigations might very however, we whether the or questions in an AVI might provide a by which to improve fairness for candidates. A of AVIs is that candidates for during the so these can be very for who are more likely to interview questions or the of For example, probes could be used to questions, or to additional information that some candidates may not recognize is expected. We also that had the call for papers and the time this was and in that time, new developments in AVIs special issue two papers that involved using AI to and deceptive interviewee behavior, and a third that focused on candidates using generative AI to to interview questions, the advances and of in AVI The however, not For example, video interview are emerging that use generative AI-based virtual to a interview This many new questions around et al. and the use of AI As Min et al. (2024) point these developments many important decisions by vendors and It is whether these can be and they from interviews by the of a human machine automated scoring approaches have also with an increased focus on fairness and validity (Hickman et al. 2024; Hickman et al. 2024; et al. and the evidence base for these approaches thus on a of vendors to the results of their We also on the implications of AVIs being The study by Arseneault and Roulin (2024) highlighted the subtle that culture can influence candidates' impression management behavior, which may have implications for candidate We also that the evidence from Basch et al. (2024) and Springle and Bourdage (2025) that AVI evaluators are of background signals of and SES may be For example, both organizational and culture could the and of types of background signals in AVIs, and these may also over time, suggesting there are many further questions about the and sources of Finally, while we the and important experimental work on AVIs, we also see a need for more field research of AVIs. field research to validity for work performance, for which evidence remains field designs also but will be for in the evidence base we are with lab in high-stakes settings where is very and we Niemitz et al. (2024) for their work In very on of high-stakes selection however, there is a of a large evidence base for the all had with the responses to AVI questions by our research we take in the that our participants to the low-stakes of an We are also by the of studies field including Benson et al.'s (2025) study in this special issue of autistic job candidates, along with other et al. et al. Stevenor et al. 2024; et al. 2024). AVI research is rapidly by the of this special some important and questions as is new forms of this We to with and the that practitioners have as they are being in the This way, these new forms could be as they are While some would to such a the in this special issue that such an would be Further, from organizational studies and be for predict the likely effects of new developments in this The would like to the of all articles submitted as of this special The authors no of
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".