Bibliographic record
Abstract
Distant Listening and Resonance Tanya E. Clement (bio) For speech recordings, sound is text—the words people speak, but also other sounds that indicate a speaking and listening context: tone and laughter, coughing and crying, bird song, car engines and horns, a baby crying, thunder clapping, gun shots, the needle dropping, the needle scratching, to name a few. Using computation to analyze many texts at once in big data sets has been called "distant reading" in Digital Humanities (Underwood). I have described "distant listening" to sound texts as using computing to "distill the many-layered four-dimensional space of the text in performance (i.e., embodied within the performance network of interpretations with the listener in time and space) into a two-dimensional script called 'code'" (Clement, "Distant Listening"). Distant listening, like distant reading, implies a lack of granular observation based on proximity in terms of space as well as a removal in terms of emotion, experience, and individual or subjective knowledge. Sound travels differently than light; what is lacking is made up for in other ways. What is too close can be too loud. What is far can be communicated loud and clear. Resonance is both an embodied, physical experience as well as a cultural hermeneutic. Specifying sound computationally is a process of discretization. Without going too far down the mathematical rabbit hole, discretization, it is safe to say, is a means of mathematically representing a continuous signal [End Page 279] through samples that indicate the whole without actually capturing it fully. Sound is air pressure variation over time. Ears turn the pressure differences into neural activations while microphones create digital sound by translating pressure differences into voltage differences. An audio signal is a sequence of mathematical abstractions that map voltage (or pressure) over time in a wave, and frequency is the number of times per second that a sound pressure wave repeats itself (McFee, "Signals"). Audio signal processing tools "cannot work directly with continuous signals," so, before being processed by a computer, the sound pressure wave must be discretized. Signal discretization includes sampling and quantization (McFee, "Digital Sampling"). The sampling process is more or less precise when more or fewer discrete samples are used to represent a signal across a period of time, but all the information is never represented. Sampling implies absence. An ontology for modeling textuality through computers requires a balance between what's computable and what is meaningful: the model should be "internally consistent, and as much as possible avoid clashes with commonsense beliefs" (Floyd and Renear). In Speech and Audio Signal Processing: Processing and Perception of Speech and Music, Ben Gold, Nelson Morgan, and Dan Ellis similarly describe this balance between meaning and matter within the history of speech transmission: If we think, for the moment, of speech as being a mode of transmitting word messages, and telegraphy as simply another mode of performing the same action, this immediately allows us to conclude that the intrinsic information rate of speech is exactly the same as that of a telegraph signal generating words at the same average rate. Speech, however, conveys emphasis, emotion, personality, etc., and we still don't know how much bandwidth is needed to transmit these kinds of information. (21) A computationally tractable model of a text, is much like a bandwidth—"a range of frequencies or wave-lengths that falls between two given limits" ("band, n.2.")—it must be explicit, consistent, and manipulable (McCarty), yet it remains always partial and inexact. Distant listening is at root a technically complex matter of fitting a mathematical abstraction of sound to a lived experience about what that sound means. When signal processing scientists talk about sound, they consider damping ratios, gain, frequencies, spectra, energy, and pitch energy and talk about how these features influence sound fidelity. When humanists talk about sound, they talk about language dynamics (tempo, [End Page 280] pitch, tone/timbre, volume, pace, laughter, silence, applause, moans, screams, dialects, changing speakers, gender, age, changing genres), environment (fan hums, car horns, chickens, train whistles, bird calls, frogs mating), and materiality (recording noises such as changing tracks, distortion, the electronic grid, and needle drops). When humanists talk about sound, they then abstract...
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".