Affective Voice Interaction and Artificial Intelligence: A Research Study on the Acoustic Features of Gender and the Emotional States of the PAD Model.

Front Psychol

Department of Digital Media Art, Design Academy, Sichuan Fine Arts Institute, Chongqing, China.

Published: May 2021

New AI products are shifting towards voice interaction, aiming to understand user emotions and provide immediate feedback.
Researchers are using deep learning to create affective acoustic models, enabling computers to predict users' emotional states based on their voice, although these models currently struggle with explaining their predictions.
The study analyzes how seven key acoustic features differ in relation to user gender and emotional states, using audio recordings and statistical analysis to establish a theoretical basis for improving AI's emotional recognition and interaction capabilities.

New types of artificial intelligence products are gradually transferring to voice interaction modes with the demand for intelligent products expanding from communication to recognizing users' emotions and instantaneous feedback. At present, affective acoustic models are constructed through deep learning and abstracted into a mathematical model, making computers learn from data and equipping them with prediction abilities. Although this method can result in accurate predictions, it has a limitation in that it lacks explanatory capability; there is an urgent need for an empirical study of the connection between acoustic features and psychology as the theoretical basis for the adjustment of model parameters. Accordingly, this study focuses on exploring the differences between seven major "acoustic features" and their physical characteristics during voice interaction with the recognition and expression of "gender" and "emotional states of the pleasure-arousal-dominance (PAD) model." In this study, 31 females and 31 males aged between 21 and 60 were invited using the stratified random sampling method for the audio recording of different emotions. Subsequently, parameter values of acoustic features were extracted using Praat voice software. Finally, parameter values were analyzed using a Two-way ANOVA, mixed-design analysis in SPSS software. Results show that gender and emotional states of the PAD model vary among seven major acoustic features. Moreover, their difference values and rankings also vary. The research conclusions lay a theoretical foundation for AI emotional voice interaction and solve deep learning's current dilemma in emotional recognition and parameter optimization of the emotional synthesis model due to the lack of explanatory power.

Download full-text PDF	Source
http://www.ncbi.nlm.nih.gov/pmc/articles/PMC8129507	PMC
http://dx.doi.org/10.3389/fpsyg.2021.664925	DOI Listing

Publication Analysis

Top Keywords

voice interaction

acoustic features

artificial intelligence

gender emotional

emotional states

states pad

pad model

parameter values

acoustic

emotional

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!

A PHP Error was encountered