MétaCan
Menu
Back to cohort

Speech Recognition System based on Wavelet Multi- Resolution Analysis using One-Dimensional CNN-LSTM Network

2024· article· en· W4403677317 on OpenAlexaboutno aff
Rekha S. Kotwal, Anju Gautam

Bibliographic record

Venuenot available
Typearticle
Languageen
FieldEngineering
TopicAdvanced Algorithms and Applications
Canadian institutionsnot available
Fundersnot available
KeywordsComputer scienceWaveletSpeech recognitionArtificial intelligencePattern recognition (psychology)Resolution (logic)Wavelet transform

Abstract

fetched live from OpenAlex

Speech Emotion Recognition (SER) is responsible for identifying the speaker’s emotions through speech and has a significant role in psychological assessment and human-computer interaction (HCI). Various time representations like Mel Spectrograms, spectrograms, as well as Mel Frequency Cepstral Coefficients (MFCCs) are widely utilized to develop SER systems. This representation uses Fast Fourier Transform (FFT) to translate the time domain signal into a frequency. However, FFT is constrained by the uncertainty principle and cannot attain great resolution in both time and frequency at the same time. In contrast, wavelets are able to offer better localization in both cases with their high resolution. Use autoencoders and combine long short term memory (LSTM) networks with one-dimensional convolutional neural networks (CNN). The scenes from the one- dimensional CNN-LSTM model are classified employing the latent space that results from the reduction of the autoencoder of the wavelet features' dimensionality. The method achieved an unweighted accuracy (UA) of 81.45% and a weighted accuracy (WA) of 81.22% when applied to the Ryerson Audiovisual Emotional Speech and Song Database (RAVDESS) dataset utilizing Monte-Carlo K-fold validation. The state- of-the-art method uses another time-frequency representation of the situation. Speech signal recognition is emerging research in human-computer interaction, and its uses incorporate human-computer interaction, the usability of virtual reality, behavioral assessment, and medical and emergency services. A novel method introduces an AI- assisted deep Stochastic convolutional neural network (DSCNN) architecture, which utilizes convolutional networks to enhance and learn important features in speech spectrograms. This model subsamples the feature maps with specific steps in the convolutional layers, bypassing the need for pooling layers and learning general features across all layers. Then, the SoftMax classifier is used for classification. Evaluation of Interactive Emotional Dyadic Motion Capture (IEMOCAP) and Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) data shows 7.85% and 4.5% accuracy and 34.5 MB reduced standard respectively, which shows the effectiveness and practicality of SER technology. Applications appear in SER due to its useful information. However, incorrect extraction rules and unclear solutions may limit the performance. To solve these problems, the psychoacoustic model inspired by speech coding introduces the information of the segmentation line to obtain a more comprehensive solution. Three new spectral characteristics— spectral flatness, spectral slope, and spectral entropy—are put forth. The hypothesis set is identified utilizing a support vector machine (SVM) classifier. Experiments indicate that this approach is more effective than other cutting-edge Fourier and multi-resolution amplitude features as well as MFCC features.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.005
Threshold uncertainty score0.012

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.000
Science and technology studies0.0000.000
Scholarly communication0.0000.001
Open science0.0010.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0030.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.035
GPT teacher head0.258
Teacher spread0.223 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2024
Admission routes1
Has abstractyes

Explore more

Same topicAdvanced Algorithms and ApplicationsFrench-language works237,207