In this work, the affective state of users in virtual learning environments is assessed/recognized in terms of continuous arousal and valence dimensions, making use of multimodal information (audio, text and video), whenever any of these modalities are available. In general, virtual learning environments where these three modalities are all the time, are not common; at some moments only the video modality is available, while in others only text or/and video and/or audio. Different approaches using feature-level fusion and decision-level fusion are proposed for multimodal recognition with missing data. Recognizing according to available modalities is studied following the ideas of dropout from neural networks and of variable input length from recurrent neural networks. This proposal is innovative because it represents emotions in the continuous space, which is not common in virtual education; and makes use of the available modalities in a virtual environment in a given moment, which is very common in virtual learning environments because the people are not speaking or writing all the time.

Download full-text PDF

Source
http://www.ncbi.nlm.nih.gov/pmc/articles/PMC8220333PMC
http://dx.doi.org/10.1016/j.heliyon.2021.e07253DOI Listing

Publication Analysis

Top Keywords

virtual learning
16
learning environments
16
affective state
8
multimodal recognition
8
neural networks
8
common virtual
8
virtual
6
analysis affective
4
state multimodal
4
recognition approaches
4

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!