MétaCan
Menu
Retour à la cohorte
Enregistrement W2387787538 · doi:10.1115/1.4033226

Affective Voice Recognition of Older Adults1

2016· article· en· W2387787538 sur OpenAlexafffundabout
Alexander Hong, Yuma Tsuboi, Goldie Nejat, B. Benhabib

Notice bibliographique

RevueJournal of Medical Devices · 2016
Typearticle
Langueen
DomainePsychology
ThématiqueEmotion and Mood Recognition
Établissements canadiensUniversity of Toronto
Organismes subventionnairesNatural Sciences and Engineering Research Council of Canada
Mots-clésLonelinessPsychologyFacial expressionAffect (linguistics)Intonation (linguistics)Body languageSocial isolationCognitionCognitive psychologySocial robotSocial relationRobotComputer scienceSocial psychologyCommunicationArtificial intelligencePsychotherapistLinguistics

Résumé

récupéré en direct d'OpenAlex

Older adults (>75 years old) may suffer from social isolation, social inactivity, or loneliness due to physical and cognitive disabilities as well as lifestyle adjustments resulting from old age [1,2]. Socially assistive robots can be used as an effective technology for the elderly to provide social interaction and cognitive assistance with activities of daily living. For example, they can support older adults with self-maintenance tasks (e.g., eating, grooming, and dressing), recreational activities (e.g., playing music and games), etc.In order to promote natural and social human–robot interaction (HRI), and provide the elderly with suitable assistance, robots would need to be equipped with emotional intelligence. For example, they would need to have the ability to consider and respond to the emotions, moods, or affect of the person with whom they are interacting [2].Older adults, including those with dementia, communicate their affective states using facial expressions, body language, and vocal intonation [3]. Our research focuses on the implementation and testing of emotion-based bidirectional interactions, to provide social and cognitive stimulation to older adults, via the intelligent socially assistive robot Brian 2.1 (Fig. 1).Our previous work with Brian 2.1 has focused on the detection of facial expressions [4] and body language [5] of the user. In this paper, we focus on the recognition and identification of affective vocal intonation of older adults as an input to determine Brian's corresponding assistive behaviors. For example, we present the development of an architecture to automatically recognize and classify affective states of older adults from vocal intonation.It has been shown that classifying affective states through voice is challenging, particularly for person-independent recognition and, furthermore, that recognition rates for older adults are lower compared to younger age groups [6]. The aging process directly affects the quality of the voice, as well as its production as a result of various physiological and anatomical changes on the vocal system [7]. For example, a valence detector was investigated in Ref. [8] using elderly voices. However, overall, with respect to automated recognition and classification of affect encompassing states of both arousal and valence during HRI scenarios, current research has not targeted the elderly population [9].Herein, we investigate the recognition and classification of the following combination of positive, neutral, and negative affective states: happy, sadness, anger, and neutral. Happiness is important to detect as for older adults it can indicate well-being, health, and longevity [10]. Sadness and anger are important to detect as they can be the signs of depression as a result of aging, for example, they are often observed in people suffering from dementia [11]. Neutral, which represents an experience of little or no noticeable feelings, is also useful to detect as a baseline for comparing other affective states.Our proposed automated vocal affect detection and classification architecture consists of three main modules: voice recognition, affect feature extraction (AFE), and affect classification (AC, Fig. 2).The VR module is responsible for capturing the audio signal of the elderly speaker and processing it into a file to be used by the AFE module in order to extract voice features from the signal (in our case, a 16-bit 11,025 Hz.wav file). This process was automated for real-time analysis by the robot. Each audio clip is 2–3 s in duration.The AFE module determines the vocal features used to classify the affective states of the elderly. In our work, we utilized the QA5 SDK Version 5.5 software by Nemesysco to identify these features. The.wav files are analyzed based on signal features such as thorns (which are local extrema in amplitude found in the second voice sample in three consecutive voice samples in a clip) and plateaus (local flatness in the voice in the clip) [12]. The output we use from the software is 18 emotion features, which are identified in the audio clip. These include content, angry, excitement, upset, energy, hesitation, embarrassment, stress, extreme state, emotion–cognition ratio, arousal factor, imagination activity, intensive thinking, concentration level, uncertainty, brain power, max amplitude volume, and voice energy.The 18 features determined are used to classify affective states. For example, within the AC module, the relationship between the affective states and the features can be identified using a machine learning technique. The following learning-based classifiers were investigated in our work: Naïve Bayes probabilistic classifier, logistic regression (LR) linear classifier, random forest (RF) decision tree, k-nearest neighbors lazy learning-based classifier, multiperceptron neural network, and nonlinear support vector machines (SVM). These techniques were considered based on their robustness to handle a wide variety of features needed to determine the affective states.In order to validate the proposed architecture, 123 audio clips from 57 older adult speakers were obtained. The participants were both males and females (≥58 years old) engaged in conversation with different intonations. The audio clips were obtained from numerous sources, including YouTube videos, talk-show interviews, news broadcasts, and the SEMAINE database [13]. Two coders were used to code the baseline ACs for each clip. The clips for which consensus was obtained between the two coders were used as the input dataset into our proposed automated vocal affect detection and classification system.A tenfold cross-validation approach was used to both train and test each classifier using the aforementioned 123 audio clips. The results are presented in Table 1. The RF decision tree and LR linear classifier provided the highest classification rate of 68.3%.The confusion matrix for the affective states for both classifiers are presented in Table 2. The highest classification rate for both classifiers was for anger (78%). The lowest classification rate was for sadness (56% for RF and 64% for LR, respectively). Sadness was challenging to recognize for all the classifiers as Nemesysco does not provide a distinctive feature to illustrate a sadness affective state.The objective of our research is to develop an emotionally intelligent socially assistive robot to assist the elderly. In this paper, we have presented an automated vocal affect recognition and classification architecture for estimating the affective states of older adults. Our results show that by using RF and LR classifiers, one can classify the affective states of happy, sadness, anger, and neutral at a rate of approximately 68%. In contrast, compared to Ref. [8], where elderly valence was classified at a rate of 55% or lower. Future work will consist of investigating and comparing our features to psychoacoustic features (i.e., loudness, tempo, contour, and sharpness), which have been directly linked to affective states [14].This work was supported by the Natural Sciences and Engineering Research Council of Canada (NSERC) and the Canada Research Chairs (CRC) Program.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,001
score de la tête « metaresearch » (Gemma)0,001
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesCharge utile insuffisante (le modèle a refusé de juger)
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Autre devis · Signal consensuel: aucune
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,982
Score d'incertitude au seuil0,993

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0010,001
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,000
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0000,000
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0080,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,029
Tête enseignante GPT0,341
Écart entre enseignants0,312 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Devis d'étudeAutre devis
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations4
Publié2016
Routes d'admission3
Résumé présentoui

Explorer davantage

Même revueJournal of Medical DevicesMême sujetEmotion and Mood RecognitionTravaux en français237 207