Sound environment and sound field reproduction using transducer arrays: Correlation between physical and perceptual evaluations
Bibliographic record
Abstract
Sound Environment Reproduction (SER) using Sound Field Reproduction (SFR) is aimed at the spatial reconstruction of a target sound field captured using a microphone array. SFR has recently gained attention for SER in industrial or engineering contexts for sound comfort or sound quality studies. The challenge is to create a reproduced sound field that first satisfies an assessment based on physical evaluation, for example to satisfy any regulation based on physical quantities. However, the reproduced sound environment should also success in perceptual evaluation. In this work, SER was applied to spatial sound field simulation in a vehicle mock-up. Both physical and perceptual evaluations were completed. Physical metrics such as frequency-dependent averaged reproduction error (both phase and magnitude) and averaged magnitude error (ignoring phase) were measured. Perceptual evaluations were based on similarity listening tests while comparing SER with an original reference (the target sound field). Perceptual evaluations were compiled as similarity scores. Correlation of similarity scores based on various physical evaluations suggests that the frequency-averaged and spatially averaged magnitude error is the physical evaluation metric that is the most correlated with the result of listening tests. This suggests that spatially reproducing the accurate frequency spectrum is the first criterion for immersing SER.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.031 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".