Enhancing Human Emotion Detection in Audio Data with Deep Neural Networks Using Cross-Dataset
Bibliographic record
Abstract
Understanding emotions from spoken language is a natural human ability, but it brings machines into great perplexity due to the complex variation in an individual’s speech. Our research focuses on working out this challenge by developing a model that makes use of deep neural networks, particularly Convolutional Neural Networks(CNN), for the recognition of emotions from speech. The goal is to be equipped with a system capable of independent emotion detection for a variety of audio datasets, adaptation to new situations, and staying accurate even in fully new cases of data. To achieve this, we combined two strong datasets: the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) and the Toronto Emotional Speech Set (TESS). The TESS dataset is particularly useful because it includes high-quality recordings from female speakers, helping to balance the gender differences often found in other datasets. This balance makes our model better at recognizing emotions across different voices. We implemented state-of-the-art techniques in audio feature extraction so that the input to our model only consists of relevant features, enabling it to extract the emotional content correctly. Our model was trained on eight emotions: neutral, calm, happy, sad, angry, fearful, disgusted, and surprised. For measuring model performance, we used the F1 score and obtained a weighted average of 0.94 on the test set. It did best on classifying the "Calm" emotion, with an accuracy of 0.97, and worst for the "Sad" emotion, with an accuracy of 0.91. Notwithstanding, our model outperforms current models in handling the huge diversity of voices in speech data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".