Envisioning AI Support during Semi-Structured Interviews Across the Expertise Spectrum
Bibliographic record
Abstract
Semi-structured interviews are a critical qualitative method in many areas, including CSCW and HCI, enabling researchers to uncover deep contextualized insights. Using this flexible method, interviewers must adapt to interviewees' responses while adhering to the protocol, necessitating strong active listening and real-time analytical skills. While recent studies have explored how AI can support researchers in qualitative analysis, to our knowledge, no research has investigated AI's role in supporting semi-structured interviews while they are underway. Taking one step toward filling this gap, we interviewed 16 researchers with a range of prior interviewing expertise. Our inductive thematic analysis reveals that interviewers expect real-time AI assistance to support research objectives and facilitate interpersonal communication, but also have concerns about its impact on long-term skill development. We discuss how semi-structured interviews differ from other problem-solving or creative human-AI collaboration contexts, highlighting the time constraints, multimodal collaboration, and the triangular dynamic among interviewers, interviewees, and AI. We also delve into how interviewers' levels of expertise affect their envisioned interviewer-AI collaboration. We then propose design challenges for future CSCW work on AI-driven assistants in interview contexts.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.126 | 0.144 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.007 | 0.011 |
| Scholarly communication | 0.006 | 0.009 |
| Open science | 0.003 | 0.009 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".