Beyond Words: Extracting Emotions from Speech with AI Techniques
Bibliographic record
Abstract
Emotions are important in daily life because they affect how individuals act, think, and react in a variety of circumstances. Thanks to current advances in Artificial Intelligence (AI), we can train machines to effectively understand and predict human emotions. AI based on; Speech Emotion Recognition (SER) helps in human-computer interaction. Devices like Smart Home Assistants can provide useful feedback thanks to the voice emotion recognition technology and the need to stay at home for a long time. This paper's main objective is to use machine learning and deep learning methods to analyse speech's emotional states. Random Forest, 1-D, and 2-D Convolutional Neural Network (CNN) are used to construct a high-quality voice emotion identification system. This method identifies seven different emotions: neutral, joyful, sad, furious, fearful, repulsed, and surprised. a standard dataset for emotion categorization used for training and testing that combines the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) and Toronto Emotional Speech Set (TESS). To extract features from our sound wave data, we worked with Mel-frequency cepstral coefficient (MFCC) and Mel Spectrogram techniques using a python audio processing library called Librosa. 85.6%, 96.6% ± 1.5% and 98.4% ± 1.43% classification accuracies were obtained using Random Forest, 1-D CNN and 2-D CNN respectively. This research concludes that, 2-D CNN classification model can be effectively used to train machines to understand and predict human emotions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".