Clarifying spectral and temporal dimensions of musical instrument timbre
Bibliographic record
Abstract
Classic studies based on multi-dimensional scaling of dissimilarity judgments, and on discrimination, for musical instrument sounds have provided converging support for the importance of relatively static, spectral cues to timbre (e.g., energy in the higher harmonics, which has been associated with perceived brightness), as well as dynamic, temporal cues (e.g., rise time, associated with perceived abruptness). Comparatively few studies have evaluated the effects of acoustic attributes on instrument identification, despite the fact that timbre recognition is an important listening goal. To assess the nature, and salience, of these cues to timbre recognition, two experiments were designed to compare discrimination and identification performance for resynthesized tones that systematically varied spectral and temporal parameters between settings for two natural instruments. Stimuli in the first experiment consisted of various combinations of spectral envelopes (manipulating the relative amplitudes of harmonics) and amplitude-vs.-time envelopes (including rise times). Listeners were most sensitive to spectral changes in both discrimination and identification tasks. Only extreme amplitude envelopes impacted performance, suggesting a binary feature based on abruptness of the attack. The second experiment sought to clarify the spectral dimension. Listener sensitivity was compared for a) modifications of spectral envelope shape via variation of formant structure and b) spectral changes that minimally impact envelope shape (using low-pass filters to match the centroids of the formant-varied envelopes). Only differences in formant structure were easily discriminated and contributed strongly to identification. Thus, it appears that listeners primarily identify timbres according to spectral envelope shape. Implications for models of instrument timbre are discussed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".