Reimagining Accents and Speech Recognition with Sociolinguistic Perception Studies and Research on Listening Subjects
Notice bibliographique
Résumé
This commentary unpacks how insights from studies of sociolinguistic perception and research on listening subjects may offer compelling ways to recognise the humanity and multiplicity of each voice, disclosing that perception is something listeners do. Below, I think with selected projects from the two strands of research to advance discussions on accent bias in technology, showing how listening is formed through practices operating within particular cultures of reception, where ‘even the most silent of listeners is an author of an emergent narrative’ (Ochs and Capps 1996: 21). I argue that the two strands enable us to reimagine all accents as loci ‘of the experience and knowledge production of the modern’ (Inoue 2003: 158), operating through particular practices of citation transcending ‘observable and […] recordable “realities”’ (Inoue 2003: 182). This in turn enables more voices ‘to be justly recognized’ (Eidsheim 2023: 143), moving beyond only listening from positions of power to a better understanding of how affordances and infrastructures amplify, mishear or silence particular human soundings. With the increased use of voice AI technologies, researchers have become interested in examining what listening practices are being translated into algorithms and how they perpetuate particular ideas about language structure and use. Recently, a group of computational linguists at Cambridge has evaluated automatic speech recognition systems (ASR), which sit at the core of many such technologies. Having dramatically improved (e.g. Jurafsky and Martin 2025), today most mainstream ASR systems recognise acoustic patterns in audio recordings and map them onto probabilities for corresponding text to transcribe human speech into writing. After testing the performance of tools provided by Google, OpenAI and Meta on corpora with ‘standard’ and ‘non-standard’ audio data in Arabic, Spanish, Bengali, Georgian, Tamil, Telugu and Tagalog, Kantharuban et al. (2024) report, however, that the tools still underperform for under-resourced dialect varieties. They argue that ASR performance ‘depends on the task and existing state of the [dialect] gap’ in ASR training datasets, with the largest predictor for having one's speech recognised being ‘linguistic proximity to well-resourced dialects’, most widely used for ASR training. Such accented listening of ASR systems makes some accents hyperaudible, perpetuating oppressive social relations. Research shows that ASR's underperformance may reinforce or even amplify existing ethnoracial or regional dialect disparities. For example, Koenecke et al. (2020) observed twice as many errors for Black as for White American English speakers, with the highest word error rate reported for Black men using most features associated with African American English. Similarly, mainstream videoconferencing and social media platforms produce twice as high an error rate for L2 speakers, who make most speakers of English, as for L1 speakers (Dubois et al. 2024). Increasingly, studies attend to the behavioural and psychological impact of ASR's underperformance for minoritised groups (Menegesha et al. 2021), with many highlighting the need to rewire algorithms for ‘just recognition’ (Eidsheim 2023). In battles against these real-life injustices in ASR, which is increasingly used in institutional contexts, perception studies remind us that human speech perception is not only about acoustic pattern recognition. Rather, human spoken word comprehension is a process embedded in ‘complex social dynamics on the one hand and rapidly occurring linguistic cues on the other’ (Campbell-Kibler 2020: 254). Thanks to their experimental design and meticulous study of factors that influence listeners’ expectations of speech, linguistic processing and memory, these projects reveal how humans use linguistic styles to contextualise the meaning of variation or integrate external information when perceiving and evaluating linguistic material. This knowledge is crucial for exposing in-built biases in ASR design. For example, Wong and Babel's (2017) study of 30 individuals self-identifying as Chinese, East Indian or White Canadian in Vancouver shows that, like ASR, humans most accurately recognise speech patterns associated with dominant groups, in this case the White Canadians. However, experience with a minority group also correlates with high accuracy and reveals consistent labelling choices of listeners, which requires ASR systems to be carefully curated and balanced for particular varieties (Bender et al. 2021) alongside advancing optimisation strategies. Based on Campbell-Kimbler (2021), Hay and Drager (2010) or D'Onforio (2019), we must, however, concede that paying attention to variation in the acoustic signal is only part of the story. In multimodal human-to-human interactions, the way audio and visual modalities combine influences perception. Campbell-Kimbler's experimental work on perception of foreign accent and speakers’ attractiveness, for example, shows that explicit instruction and the task performed have a strong influence on the role face and voice information play in social perception. After presenting 1034 US-based participants, mostly self-identifying as female, White and in mid-twenties, with speech stimuli of single words pronounced with different degrees of accentedness together with 85 still images of male faces with uniform bodily postures, different background colours and shirt styles, Campbell-Kimbler shows that perception is shaped through different types of information, not always deliberately controlled and reviewed. When discussing accentedness, D'Onforio further notes that visual styles provide rich information that may influence linguistic memory as expectations of accentedness are not mapped directly onto racialised categories based on phenotypical information alone. This is evident when 153 L1 American English-speaking participants, with different degrees of familiarity with Korea, link the visual stimuli with a White man and the same Korean male actor representing different personae to audio samples with a recorded passage produced either by an L2 English speaker with L1 Korean or by an L1 American English speaker. With one Korean persona rated as ‘warmer, more likeable’, ‘more American, more casual’ and more likely to be ‘from the US’ than the other, the results show disparate recall accuracy rates, with expectations of the former patterning more closely with those associated with a White man. When considering ASR design, most models do not integrate multiple information they receive in dynamic processes in a similar way, which is itself not a neutral calibration of digital-audio information. Recognising only some person types, ASR systems also propagate them as most representative of ethnoracialised groups, building into particular discourses of authenticity. While studies of sociolinguistic perception uncover how ASR systems listen with an accent built through such a selective use of images of others and cues, their experimental design momentarily freezes social relations and mostly stresses the impossibility of neutral listening. It is then research on listening subjects (e.g. Rosa and Flores 2017; Pak 2023) that pushes us to deal with this impossibility: to investigate what happens to this multimodality of human listening when human voices get rerouted through ASR infrastructures, become embedded in particular habits and contexts of use and how and for whom they recreate realities of those who heard particular voices and understood accents in particular ways making the systems recognise in-built accents and voices operating in their ‘acoustic shadows’ in different ways (Eidsheim 2023). Sociolinguistic perception studies also often focus on isolated variables and manipulate them to test causal relationships through perception experiments, acoustic manipulation of stimuli or forced-choice judgements, treating the listener as revealing cognitive processes and social indexing practices. In contrast, research on listening subjects employs long-term participant observations, situated interviewing techniques or work with archives to argue that the listener is also a social actor whose listening is shaped by the lived experience and remains selective and ideological. By doing so, they highlight that ASR adapts certain assumptions into algorithms, erasing the reality in which perception is ‘an effect of a regime of social power’ at a particular time and place and ‘never a natural or unmediated phenomenon’ (Inoue 2003: 157). Therefore, research on listening subjects compels us to interrogate the conditions of this cultural practice: who listens, who gets heard, when and how, also outside of carefully designed research experiments, and what counts as audible and inaudible given the institutional history of registering particular signs, media infrastructures, racialisation processes or colonial histories. By doing so, it makes us acutely aware that not all others can ‘constitute themselves’ and ‘speak for themselves’ as those in dominant positions, which in turn may help better understand how ASR structures of sociality naturalise shades of whiteness and interconnected systems of oppression, and make not properly sounding humans develop new strategies to have their humanity recognised. In order to understand such ‘processes of contingent, collaborative and emergent self-making’ (Smalls 2018: 360) in human–machine interactions, these projects oblige us to look at the ways in which ‘[s]poken interactions with voice-AI influence human speech patterns in socially-meaningful ways’ and ‘users’ social characteristics […] shape their attitudes and accommodative behaviour towards machines' (Zellou and Holliday 2024: 5) at various intersections of society. They mandate further research examining how particular epistemes manifest across groups, what naturalising events and means enable audible performances and what human effort is involved in the construction of the self ‘across individual discursive fragments’ (Smalls 2018: 361). Building on Smalls's ethnographic work on antiblackness and discursive violence in a high school in Pennsylvania, it could be argued that the most pressing question that emerges is not about ASR's accuracy, but how the meanings of ASR representations are made through interdiscursive chaining of events and relevant practices of citation through which they are ‘made, recorded, and legitimated through linguistic and other means of circulation’ (Smalls 2018: 361). This way, we may better grasp barriers created through the use of voice AI technologies in relation to socio-spatial relations in which groups are made subordinated or subaltern, as well as the ‘ontological violence that hinder[s one's] ability to freely create and sincerely express whole, black selves’ (Smalls 2018: 377) in different contexts of use and through changing experiences with technology. It is hence necessary to further investigate how the emergence of communicative routines with machines is amplified, reinforced or hindered through sensory transformations enabled by voice AIs at the margins and how inaudible interlocutors make their language perceptible (Edwards 2018). Building on my own work (Kozminska 2024) examining modes of transformation among moving transnational actors speaking not just English, but also Polish in which I draw on Rancière's (2013) concept of the ‘distribution of the sensible’, it is also vital to disclose such an emerging ‘system of self-evident facts of sense perception’ in human–machine interactions, what individuals do with nascent possibilities and how they establish their relation with otherness and the self in situated events. Following Edwards's work with DeafBlind people at Gallaudet University, it must, at the same time, be remembered that adaptations to sensory channels depend not only on the relation between interlocutors and ways they are woven into other relations but also on contrasts between channels and their reroutings through and with the environment. Working intersectionally on such reroutings may enable us to elucidate how ‘our environment doesn't just supply us with the specifics to go with our schemas, [but i]t anticipates us’ (Edwards 2018: 277) and how, through the ASR infrastructure, sensory orientation is shaped or reinforced, making us perceive affordances in particular ways. This commentary thus echoes Eidsheim's call to ‘listen to listening’ and emerging ASR cultures of reception as a way to reveal how human and human-through-machine-propagated fantasies of just recognition built on ideals of accented listening are shaped by past histories of contact (Ahmed 2014). The shared interest of the two research strands in the ways in which hearing is socially (re)shaped urges us to move away from voices imagined as static and monolithic to those recognised in ‘relationship to oneself and to the multiplicity of histories and communities’ (Eidsheim 2023: 144) in flux. It compels us to pay close attention to generated distinctions, their embedding in past associations and emotional responses, reminding us that the politics of healing is not only enacted in the moments of hearing and that making more voices audible is always conditioned by legible configurations of time–space–personhood. Following Kantharuban et al. (2024), the two strands combined therefore bring into sharp relief the fact that when being recognised, the issues a ‘non-standard’ ASR user of under-resourced languages may face may be difficult for a ‘standard’ user of well-resourced languages such as English to imagine. However, they not only do so but also attempt to convince the former that they too can reimagine speech recognition.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,002 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».