In the Eye of the Transcriber: Four Column Analysis Structure for Qualitative Research With Audiovisual Data
Bibliographic record
Abstract
For the past thirty years, qualitative psychology researchers have focused on the study of written or spoken word while relegating the study of visual communications to children or those deemed unable to speak (Reavey, 2021, p. 2). Thus, the discipline now pays much attention to the collection and analysis of spoken or written words but significantly less to visual and auditory expressions of experience (Reavey, 2021, p. 3), and most transcription methods for psychology researchers are those designed for interviews that only capture the spoken word. However, these transcription methods have yet to account for the current context of ubiquitous, technologically mediated interactions. People from diverse groups use social media platforms such as YouTube and TikTok to interact using speech, audio and video. While they offer rich data for qualitative psychology researchers, the tools to capture such multimodal expressions are still in early stages of development within the discipline (Marshall et al., 2021). In this article, we present a transcription structure that allows for the recording of both speech and visual elements in audiovisual content. Inspired by methods from communications and visual anthropology, the Four Column Analysis Structure (or, FoCAS) allows for the simultaneous analysis of both audio and visual data by allowing for the transcription of four dimensions: (1) timestamp, (2) setting, (3) scene, and (4) audio. Based on its application in two completed studies and one study in progress, we describe the development of the FoCAS, how to set it up, transcription conventions, and how to analyze qualitative data using all four columns. We additionally discuss sampling considerations and the advantages and disadvantages of the structure. By expanding the amount of meaningful data that can be captured by qualitative transcription, we hope the FoCAS can be used to create more multidimensional, rigorous analyses of audiovisual data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.022 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".