Assessing Basic Emotion via Machine Learning: Comparative Analysis of Number of Basic Emotions and Algorithms
Bibliographic record
Abstract
This paper explores the use of machine learning (ML) methods to identify "clusters" of basic emotions based on pleasure, arousal, and dominance (PAD). The data was obtained from the Dataset for Emotion Analysis using Physiological Signals (DEAP), data collected within the Building and Designing Assistive Technology (BDAT) Lab using the International Affective Picture System (IAPS), and the scores of PAD from the IAPS. The objective is to develop an algorithm that maps a PAD score to clusters that express emotions, e.g., sadness or happiness. The elbow method was used to determine the optimal number of clusters (4-8), and nine different ML algorithms were compared. Decision Trees, polynomial support vector machines (SVMs) and linear SVMs provided accurate results. The Decision Tree demonstrated efficiency, during both testing and validation, in identifying the same clusters when analyzing both the DEAP and IAPS datasets. The dataset included limited data for each emotion creating the possibility of overfitting. However, when evaluating the results relative to previous research, the results added to the understanding of the nuances of emotion self-reporting and modelling.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.007 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".