MétaCan
Menu
Back to cohort
Record W4385591415 · doi:10.1353/esc.2020.a903548

Distant Listening and Resonance

2020· article· en· W4385591415 on OpenAlexvenueno aff
Tanya Clement

Bibliographic record

VenueEnglish studies in Canada · 2020
Typearticle
Languageen
FieldComputer Science
TopicMusic and Audio Processing
Canadian institutionsnot available
Fundersnot available
KeywordsActive listeningReading (process)Context (archaeology)Embodied cognitionLaughterSpace (punctuation)Tone (literature)Computer scienceAcousticsPsychologyCommunicationLinguisticsHistoryArtificial intelligencePhysicsSocial psychologyPhilosophy

Abstract

fetched live from OpenAlex

Distant Listening and Resonance Tanya E. Clement (bio) For speech recordings, sound is text—the words people speak, but also other sounds that indicate a speaking and listening context: tone and laughter, coughing and crying, bird song, car engines and horns, a baby crying, thunder clapping, gun shots, the needle dropping, the needle scratching, to name a few. Using computation to analyze many texts at once in big data sets has been called "distant reading" in Digital Humanities (Underwood). I have described "distant listening" to sound texts as using computing to "distill the many-layered four-dimensional space of the text in performance (i.e., embodied within the performance network of interpretations with the listener in time and space) into a two-dimensional script called 'code'" (Clement, "Distant Listening"). Distant listening, like distant reading, implies a lack of granular observation based on proximity in terms of space as well as a removal in terms of emotion, experience, and individual or subjective knowledge. Sound travels differently than light; what is lacking is made up for in other ways. What is too close can be too loud. What is far can be communicated loud and clear. Resonance is both an embodied, physical experience as well as a cultural hermeneutic. Specifying sound computationally is a process of discretization. Without going too far down the mathematical rabbit hole, discretization, it is safe to say, is a means of mathematically representing a continuous signal [End Page 279] through samples that indicate the whole without actually capturing it fully. Sound is air pressure variation over time. Ears turn the pressure differences into neural activations while microphones create digital sound by translating pressure differences into voltage differences. An audio signal is a sequence of mathematical abstractions that map voltage (or pressure) over time in a wave, and frequency is the number of times per second that a sound pressure wave repeats itself (McFee, "Signals"). Audio signal processing tools "cannot work directly with continuous signals," so, before being processed by a computer, the sound pressure wave must be discretized. Signal discretization includes sampling and quantization (McFee, "Digital Sampling"). The sampling process is more or less precise when more or fewer discrete samples are used to represent a signal across a period of time, but all the information is never represented. Sampling implies absence. An ontology for modeling textuality through computers requires a balance between what's computable and what is meaningful: the model should be "internally consistent, and as much as possible avoid clashes with commonsense beliefs" (Floyd and Renear). In Speech and Audio Signal Processing: Processing and Perception of Speech and Music, Ben Gold, Nelson Morgan, and Dan Ellis similarly describe this balance between meaning and matter within the history of speech transmission: If we think, for the moment, of speech as being a mode of transmitting word messages, and telegraphy as simply another mode of performing the same action, this immediately allows us to conclude that the intrinsic information rate of speech is exactly the same as that of a telegraph signal generating words at the same average rate. Speech, however, conveys emphasis, emotion, personality, etc., and we still don't know how much bandwidth is needed to transmit these kinds of information. (21) A computationally tractable model of a text, is much like a bandwidth—"a range of frequencies or wave-lengths that falls between two given limits" ("band, n.2.")—it must be explicit, consistent, and manipulable (McCarty), yet it remains always partial and inexact. Distant listening is at root a technically complex matter of fitting a mathematical abstraction of sound to a lived experience about what that sound means. When signal processing scientists talk about sound, they consider damping ratios, gain, frequencies, spectra, energy, and pitch energy and talk about how these features influence sound fidelity. When humanists talk about sound, they talk about language dynamics (tempo, [End Page 280] pitch, tone/timbre, volume, pace, laughter, silence, applause, moans, screams, dialects, changing speakers, gender, age, changing genres), environment (fan hums, car horns, chickens, train whistles, bird calls, frogs mating), and materiality (recording noises such as changing tracks, distortion, the electronic grid, and needle drops). When humanists talk about sound, they then abstract...

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.578
Threshold uncertainty score0.917

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.032
GPT teacher head0.242
Teacher spread0.211 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2020
Admission routes1
Has abstractyes

Explore more

Same venueEnglish studies in CanadaSame topicMusic and Audio ProcessingFrench-language works237,207