MétaCan
Menu
Back to cohort
Record W4415567796 · doi:10.1111/josl.12720

Reimagining Accents and Speech Recognition with Sociolinguistic Perception Studies and Research on Listening Subjects

2025· article· en· W4415567796 on OpenAlexaboutno aff
Kinga Koźmińska

Bibliographic record

VenueJournal of Sociolinguistics · 2025
Typearticle
Languageen
FieldArts and Humanities
TopicLanguage, Discourse, Communication Strategies
Canadian institutionsnot available
Fundersnot available
KeywordsActive listeningStress (linguistics)PerceptionSilenceSpeech perceptionMainstreamAffordance

Abstract

fetched live from OpenAlex

This commentary unpacks how insights from studies of sociolinguistic perception and research on listening subjects may offer compelling ways to recognise the humanity and multiplicity of each voice, disclosing that perception is something listeners do. Below, I think with selected projects from the two strands of research to advance discussions on accent bias in technology, showing how listening is formed through practices operating within particular cultures of reception, where ‘even the most silent of listeners is an author of an emergent narrative’ (Ochs and Capps 1996: 21). I argue that the two strands enable us to reimagine all accents as loci ‘of the experience and knowledge production of the modern’ (Inoue 2003: 158), operating through particular practices of citation transcending ‘observable and […] recordable “realities”’ (Inoue 2003: 182). This in turn enables more voices ‘to be justly recognized’ (Eidsheim 2023: 143), moving beyond only listening from positions of power to a better understanding of how affordances and infrastructures amplify, mishear or silence particular human soundings. With the increased use of voice AI technologies, researchers have become interested in examining what listening practices are being translated into algorithms and how they perpetuate particular ideas about language structure and use. Recently, a group of computational linguists at Cambridge has evaluated automatic speech recognition systems (ASR), which sit at the core of many such technologies. Having dramatically improved (e.g. Jurafsky and Martin 2025), today most mainstream ASR systems recognise acoustic patterns in audio recordings and map them onto probabilities for corresponding text to transcribe human speech into writing. After testing the performance of tools provided by Google, OpenAI and Meta on corpora with ‘standard’ and ‘non-standard’ audio data in Arabic, Spanish, Bengali, Georgian, Tamil, Telugu and Tagalog, Kantharuban et al. (2024) report, however, that the tools still underperform for under-resourced dialect varieties. They argue that ASR performance ‘depends on the task and existing state of the [dialect] gap’ in ASR training datasets, with the largest predictor for having one's speech recognised being ‘linguistic proximity to well-resourced dialects’, most widely used for ASR training. Such accented listening of ASR systems makes some accents hyperaudible, perpetuating oppressive social relations. Research shows that ASR's underperformance may reinforce or even amplify existing ethnoracial or regional dialect disparities. For example, Koenecke et al. (2020) observed twice as many errors for Black as for White American English speakers, with the highest word error rate reported for Black men using most features associated with African American English. Similarly, mainstream videoconferencing and social media platforms produce twice as high an error rate for L2 speakers, who make most speakers of English, as for L1 speakers (Dubois et al. 2024). Increasingly, studies attend to the behavioural and psychological impact of ASR's underperformance for minoritised groups (Menegesha et al. 2021), with many highlighting the need to rewire algorithms for ‘just recognition’ (Eidsheim 2023). In battles against these real-life injustices in ASR, which is increasingly used in institutional contexts, perception studies remind us that human speech perception is not only about acoustic pattern recognition. Rather, human spoken word comprehension is a process embedded in ‘complex social dynamics on the one hand and rapidly occurring linguistic cues on the other’ (Campbell-Kibler 2020: 254). Thanks to their experimental design and meticulous study of factors that influence listeners’ expectations of speech, linguistic processing and memory, these projects reveal how humans use linguistic styles to contextualise the meaning of variation or integrate external information when perceiving and evaluating linguistic material. This knowledge is crucial for exposing in-built biases in ASR design. For example, Wong and Babel's (2017) study of 30 individuals self-identifying as Chinese, East Indian or White Canadian in Vancouver shows that, like ASR, humans most accurately recognise speech patterns associated with dominant groups, in this case the White Canadians. However, experience with a minority group also correlates with high accuracy and reveals consistent labelling choices of listeners, which requires ASR systems to be carefully curated and balanced for particular varieties (Bender et al. 2021) alongside advancing optimisation strategies. Based on Campbell-Kimbler (2021), Hay and Drager (2010) or D'Onforio (2019), we must, however, concede that paying attention to variation in the acoustic signal is only part of the story. In multimodal human-to-human interactions, the way audio and visual modalities combine influences perception. Campbell-Kimbler's experimental work on perception of foreign accent and speakers’ attractiveness, for example, shows that explicit instruction and the task performed have a strong influence on the role face and voice information play in social perception. After presenting 1034 US-based participants, mostly self-identifying as female, White and in mid-twenties, with speech stimuli of single words pronounced with different degrees of accentedness together with 85 still images of male faces with uniform bodily postures, different background colours and shirt styles, Campbell-Kimbler shows that perception is shaped through different types of information, not always deliberately controlled and reviewed. When discussing accentedness, D'Onforio further notes that visual styles provide rich information that may influence linguistic memory as expectations of accentedness are not mapped directly onto racialised categories based on phenotypical information alone. This is evident when 153 L1 American English-speaking participants, with different degrees of familiarity with Korea, link the visual stimuli with a White man and the same Korean male actor representing different personae to audio samples with a recorded passage produced either by an L2 English speaker with L1 Korean or by an L1 American English speaker. With one Korean persona rated as ‘warmer, more likeable’, ‘more American, more casual’ and more likely to be ‘from the US’ than the other, the results show disparate recall accuracy rates, with expectations of the former patterning more closely with those associated with a White man. When considering ASR design, most models do not integrate multiple information they receive in dynamic processes in a similar way, which is itself not a neutral calibration of digital-audio information. Recognising only some person types, ASR systems also propagate them as most representative of ethnoracialised groups, building into particular discourses of authenticity. While studies of sociolinguistic perception uncover how ASR systems listen with an accent built through such a selective use of images of others and cues, their experimental design momentarily freezes social relations and mostly stresses the impossibility of neutral listening. It is then research on listening subjects (e.g. Rosa and Flores 2017; Pak 2023) that pushes us to deal with this impossibility: to investigate what happens to this multimodality of human listening when human voices get rerouted through ASR infrastructures, become embedded in particular habits and contexts of use and how and for whom they recreate realities of those who heard particular voices and understood accents in particular ways making the systems recognise in-built accents and voices operating in their ‘acoustic shadows’ in different ways (Eidsheim 2023). Sociolinguistic perception studies also often focus on isolated variables and manipulate them to test causal relationships through perception experiments, acoustic manipulation of stimuli or forced-choice judgements, treating the listener as revealing cognitive processes and social indexing practices. In contrast, research on listening subjects employs long-term participant observations, situated interviewing techniques or work with archives to argue that the listener is also a social actor whose listening is shaped by the lived experience and remains selective and ideological. By doing so, they highlight that ASR adapts certain assumptions into algorithms, erasing the reality in which perception is ‘an effect of a regime of social power’ at a particular time and place and ‘never a natural or unmediated phenomenon’ (Inoue 2003: 157). Therefore, research on listening subjects compels us to interrogate the conditions of this cultural practice: who listens, who gets heard, when and how, also outside of carefully designed research experiments, and what counts as audible and inaudible given the institutional history of registering particular signs, media infrastructures, racialisation processes or colonial histories. By doing so, it makes us acutely aware that not all others can ‘constitute themselves’ and ‘speak for themselves’ as those in dominant positions, which in turn may help better understand how ASR structures of sociality naturalise shades of whiteness and interconnected systems of oppression, and make not properly sounding humans develop new strategies to have their humanity recognised. In order to understand such ‘processes of contingent, collaborative and emergent self-making’ (Smalls 2018: 360) in human–machine interactions, these projects oblige us to look at the ways in which ‘[s]poken interactions with voice-AI influence human speech patterns in socially-meaningful ways’ and ‘users’ social characteristics […] shape their attitudes and accommodative behaviour towards machines' (Zellou and Holliday 2024: 5) at various intersections of society. They mandate further research examining how particular epistemes manifest across groups, what naturalising events and means enable audible performances and what human effort is involved in the construction of the self ‘across individual discursive fragments’ (Smalls 2018: 361). Building on Smalls's ethnographic work on antiblackness and discursive violence in a high school in Pennsylvania, it could be argued that the most pressing question that emerges is not about ASR's accuracy, but how the meanings of ASR representations are made through interdiscursive chaining of events and relevant practices of citation through which they are ‘made, recorded, and legitimated through linguistic and other means of circulation’ (Smalls 2018: 361). This way, we may better grasp barriers created through the use of voice AI technologies in relation to socio-spatial relations in which groups are made subordinated or subaltern, as well as the ‘ontological violence that hinder[s one's] ability to freely create and sincerely express whole, black selves’ (Smalls 2018: 377) in different contexts of use and through changing experiences with technology. It is hence necessary to further investigate how the emergence of communicative routines with machines is amplified, reinforced or hindered through sensory transformations enabled by voice AIs at the margins and how inaudible interlocutors make their language perceptible (Edwards 2018). Building on my own work (Kozminska 2024) examining modes of transformation among moving transnational actors speaking not just English, but also Polish in which I draw on Rancière's (2013) concept of the ‘distribution of the sensible’, it is also vital to disclose such an emerging ‘system of self-evident facts of sense perception’ in human–machine interactions, what individuals do with nascent possibilities and how they establish their relation with otherness and the self in situated events. Following Edwards's work with DeafBlind people at Gallaudet University, it must, at the same time, be remembered that adaptations to sensory channels depend not only on the relation between interlocutors and ways they are woven into other relations but also on contrasts between channels and their reroutings through and with the environment. Working intersectionally on such reroutings may enable us to elucidate how ‘our environment doesn't just supply us with the specifics to go with our schemas, [but i]t anticipates us’ (Edwards 2018: 277) and how, through the ASR infrastructure, sensory orientation is shaped or reinforced, making us perceive affordances in particular ways. This commentary thus echoes Eidsheim's call to ‘listen to listening’ and emerging ASR cultures of reception as a way to reveal how human and human-through-machine-propagated fantasies of just recognition built on ideals of accented listening are shaped by past histories of contact (Ahmed 2014). The shared interest of the two research strands in the ways in which hearing is socially (re)shaped urges us to move away from voices imagined as static and monolithic to those recognised in ‘relationship to oneself and to the multiplicity of histories and communities’ (Eidsheim 2023: 144) in flux. It compels us to pay close attention to generated distinctions, their embedding in past associations and emotional responses, reminding us that the politics of healing is not only enacted in the moments of hearing and that making more voices audible is always conditioned by legible configurations of time–space–personhood. Following Kantharuban et al. (2024), the two strands combined therefore bring into sharp relief the fact that when being recognised, the issues a ‘non-standard’ ASR user of under-resourced languages may face may be difficult for a ‘standard’ user of well-resourced languages such as English to imagine. However, they not only do so but also attempt to convince the former that they too can reimagine speech recognition.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.002
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Qualitative · Consensus signal: Qualitative
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.086
Threshold uncertainty score0.584

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.002
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0010.001
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.203
GPT teacher head0.444
Teacher spread0.240 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designQualitative
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJournal of SociolinguisticsSame topicLanguage, Discourse, Communication StrategiesFrench-language works237,207