Branch-MFA-TDNN: A Parallel Branch Speaker Verification Model for Voice IoT
Bibliographic record
Abstract
The security of voice control in the Voice Internet of Things (Voice IoT) heavily relies on the fast and accurate authentication of the command issuer. In this work, we focus on the critical application scenario of Voice IoT in underground coal mines, where voice commands typically last 4–10 seconds. Speech in this scenario typically consists of short, imperative utterances and faces challenges from environmental noise and device heterogeneity. The limitations of traditional speaker verification models in temporal modeling restrict their performance in such scenarios. To address this, this paper proposes a three-dimensional attention module (Branch-MFA) designed for Voice IoT. This module employs a dual-parallel branch architecture: the MFA branch is responsible for extracting attention in the frequency and channel dimensions, and its multi-scale nature enables it to effectively focus on speaker-discriminative frequency bands that remain stable under noise and different collection devices, thereby enhancing the model’s environmental robustness; the GLTA branch, through its innovative grouped variable-length attention mechanism, specifically models the temporal structure of these short voice commands, addressing the challenge of sparse temporal information in short utterances. By integrating the dual-branch outputs through a fusion module, we construct the Branch-MFA-TDNN model. Experiments on the Cn-Celeb dataset show that this model significantly outperforms baseline models in short-utterance verification tasks, particularly for the challenging 4–10 second duration relevant to mine communications, providing an identity authentication solution for Voice IoT that combines high security and real-time performance. We have also released the code<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> for future comparison.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".