MétaCan
Menu
Back to cohort
Record W2950129007 · doi:10.82308/31575

Auditory-based noise-robust audio classification algorithms

2008· article· en· W2950129007 on OpenAlexfundno aff
Wei Chu

Bibliographic record

VenueeScholarship@McGill (McGill) · 2008
Typearticle
Languageen
FieldComputer Science
TopicMusic and Audio Processing
Canadian institutionsnot available
FundersNatural Sciences and Engineering Research Council of Canada
KeywordsComputer scienceNoise (video)Speech recognitionFrequency domainAudio analyzerComputational complexity theoryGaussian noiseFast Fourier transformAlgorithmPattern recognition (psychology)Artificial intelligenceSpeech codingAudio signal processingAudio signal

Abstract

fetched live from OpenAlex

The past decade has seen extensive research on audio classification algorithms which playa key role in multimedia applications, such as the retrieval of audio information from an audio or audiovisual database. However, the effect of background noise on the performance of classification has not been widely investigated. Motivated by the noise-suppression property of the early auditory (EA) model presented by Wang and Shamma, we seek in this thesis to further investigate this property and to develop improved algorithms for audio classification in the presence of background noise. With respect to the limitation of the original analysis, a better yet mathematically tractable approximation approach is first proposed wherein the Gaussian cumulative distribution function is used to derive a new closed-form expression of the auditory spectrum at the output of the EA model, and to conduct relevant analysis. Considering the computational complexity of the original EA model, a simplified auditory spectrum is proposed, wherein the underlying analysis naturally leads to frequency-domain approximation for further reduction in the computational complexity. Based on this time-domain approximation, a simplified FFT-based spectrum is proposed wherein a local spectral self-normalization is implemented. An improved implementation of this spectrum is further proposed to calculate a so-called FFT-based auditory spectrum, which allows more flexibility in the extraction of noise-robust audio features. To evaluate the performance of the above FFT-based spectra, speech/music/noise and noise/non-noise classification experiments are conducted wherein a support vector machine algorithm (SVMstruct) and a decision tree learning algorithm (C4.5) are used as the classifiers. Several features are used for the classification, including the conventional mel-frequency cepstral coefficient (MFCC) features as well as DCT-based and spectral features derived from the proposed FFT-based spectra. Compared to the conventional features, the auditory-related features show more robust performance in mismatched test cases. Test results also indicate that the performance of the proposed FFT-based auditory spectrum is slightly better than that of the original auditory spectrum, while its computational complexity is reduced by an order of magnitude. Finally, to further explore the proposed FFT-based auditory spectrum from a practical audio classification perspective, a floating-point DSP implementation is developed and optimized on the TMS320C6713 DSP Starter Kit (DSK) from Texas Instruments.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow), Science and technology studies, Insufficient payload (model declined to judge)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Other design · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.709
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.000
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.001
Science and technology studies0.0020.000
Scholarly communication0.0000.002
Open science0.0020.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.048
GPT teacher head0.235
Teacher spread0.187 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designOther design
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2008
Admission routes1
Has abstractyes

Explore more

Same venueeScholarship@McGill (McGill)Same topicMusic and Audio ProcessingFrench-language works237,207