Two-Step Microphone Array Fusion Algorithm for Enhanced Indoor Sound Source Localization
Bibliographic record
Abstract
This paper introduces a novel two-step algorithm for microphone array fusion to enhance Sound Source Localization (SSL) in indoor reverberant environments. The proposed method intelligently selects Angle of Arrival (AoA) estimates to reduce localization errors while maintaining computational efficiency. Through simulation analysis using both simulated and real Room Impulse Responses (RIRs), we identify that AoA accuracy varies depending on the sound source location, leading to unreliable estimates from certain microphone arrays. To address this, we propose a method to exclude these unreliable AoAs from SSL, improving overall localization performance.To further evaluate the effectiveness of the proposed approach, we compare it to a deep learning-based SSL method, where a Deep Neural Network (DNN) predicts source locations based on estimated AoAs. First, we compare our method to an approach that uses all available AoAs without selection, demonstrating that the two-step algorithm reduces Mean Absolute Error (MAE) by up to 50%. Next, we compare our method with the DNN-based approach, which achieves a 6.6% lower MAE while having higher 25th and 75th percentile values, and is computationally more complex and requires extensive training data. These results highlight the ability of the two-step method to efficiently determine which AoAs to use in order to maintain more accurate SSL. These findings emphasize the practical value of the proposed method in improving SSL accuracy in challenging acoustic conditions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".