Affective Voice Recognition of Older Adults1
Bibliographic record
Abstract
Older adults (>75 years old) may suffer from social isolation, social inactivity, or loneliness due to physical and cognitive disabilities as well as lifestyle adjustments resulting from old age [1,2]. Socially assistive robots can be used as an effective technology for the elderly to provide social interaction and cognitive assistance with activities of daily living. For example, they can support older adults with self-maintenance tasks (e.g., eating, grooming, and dressing), recreational activities (e.g., playing music and games), etc.In order to promote natural and social human–robot interaction (HRI), and provide the elderly with suitable assistance, robots would need to be equipped with emotional intelligence. For example, they would need to have the ability to consider and respond to the emotions, moods, or affect of the person with whom they are interacting [2].Older adults, including those with dementia, communicate their affective states using facial expressions, body language, and vocal intonation [3]. Our research focuses on the implementation and testing of emotion-based bidirectional interactions, to provide social and cognitive stimulation to older adults, via the intelligent socially assistive robot Brian 2.1 (Fig. 1).Our previous work with Brian 2.1 has focused on the detection of facial expressions [4] and body language [5] of the user. In this paper, we focus on the recognition and identification of affective vocal intonation of older adults as an input to determine Brian's corresponding assistive behaviors. For example, we present the development of an architecture to automatically recognize and classify affective states of older adults from vocal intonation.It has been shown that classifying affective states through voice is challenging, particularly for person-independent recognition and, furthermore, that recognition rates for older adults are lower compared to younger age groups [6]. The aging process directly affects the quality of the voice, as well as its production as a result of various physiological and anatomical changes on the vocal system [7]. For example, a valence detector was investigated in Ref. [8] using elderly voices. However, overall, with respect to automated recognition and classification of affect encompassing states of both arousal and valence during HRI scenarios, current research has not targeted the elderly population [9].Herein, we investigate the recognition and classification of the following combination of positive, neutral, and negative affective states: happy, sadness, anger, and neutral. Happiness is important to detect as for older adults it can indicate well-being, health, and longevity [10]. Sadness and anger are important to detect as they can be the signs of depression as a result of aging, for example, they are often observed in people suffering from dementia [11]. Neutral, which represents an experience of little or no noticeable feelings, is also useful to detect as a baseline for comparing other affective states.Our proposed automated vocal affect detection and classification architecture consists of three main modules: voice recognition, affect feature extraction (AFE), and affect classification (AC, Fig. 2).The VR module is responsible for capturing the audio signal of the elderly speaker and processing it into a file to be used by the AFE module in order to extract voice features from the signal (in our case, a 16-bit 11,025 Hz.wav file). This process was automated for real-time analysis by the robot. Each audio clip is 2–3 s in duration.The AFE module determines the vocal features used to classify the affective states of the elderly. In our work, we utilized the QA5 SDK Version 5.5 software by Nemesysco to identify these features. The.wav files are analyzed based on signal features such as thorns (which are local extrema in amplitude found in the second voice sample in three consecutive voice samples in a clip) and plateaus (local flatness in the voice in the clip) [12]. The output we use from the software is 18 emotion features, which are identified in the audio clip. These include content, angry, excitement, upset, energy, hesitation, embarrassment, stress, extreme state, emotion–cognition ratio, arousal factor, imagination activity, intensive thinking, concentration level, uncertainty, brain power, max amplitude volume, and voice energy.The 18 features determined are used to classify affective states. For example, within the AC module, the relationship between the affective states and the features can be identified using a machine learning technique. The following learning-based classifiers were investigated in our work: Naïve Bayes probabilistic classifier, logistic regression (LR) linear classifier, random forest (RF) decision tree, k-nearest neighbors lazy learning-based classifier, multiperceptron neural network, and nonlinear support vector machines (SVM). These techniques were considered based on their robustness to handle a wide variety of features needed to determine the affective states.In order to validate the proposed architecture, 123 audio clips from 57 older adult speakers were obtained. The participants were both males and females (≥58 years old) engaged in conversation with different intonations. The audio clips were obtained from numerous sources, including YouTube videos, talk-show interviews, news broadcasts, and the SEMAINE database [13]. Two coders were used to code the baseline ACs for each clip. The clips for which consensus was obtained between the two coders were used as the input dataset into our proposed automated vocal affect detection and classification system.A tenfold cross-validation approach was used to both train and test each classifier using the aforementioned 123 audio clips. The results are presented in Table 1. The RF decision tree and LR linear classifier provided the highest classification rate of 68.3%.The confusion matrix for the affective states for both classifiers are presented in Table 2. The highest classification rate for both classifiers was for anger (78%). The lowest classification rate was for sadness (56% for RF and 64% for LR, respectively). Sadness was challenging to recognize for all the classifiers as Nemesysco does not provide a distinctive feature to illustrate a sadness affective state.The objective of our research is to develop an emotionally intelligent socially assistive robot to assist the elderly. In this paper, we have presented an automated vocal affect recognition and classification architecture for estimating the affective states of older adults. Our results show that by using RF and LR classifiers, one can classify the affective states of happy, sadness, anger, and neutral at a rate of approximately 68%. In contrast, compared to Ref. [8], where elderly valence was classified at a rate of 55% or lower. Future work will consist of investigating and comparing our features to psychoacoustic features (i.e., loudness, tempo, contour, and sharpness), which have been directly linked to affective states [14].This work was supported by the Natural Sciences and Engineering Research Council of Canada (NSERC) and the Canada Research Chairs (CRC) Program.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.008 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".