Semi-supervised Multi-modal Emotion Recognition with Cross-Modal Distribution Matching

Liang, Jingjun; Li, Ruichen; Jin, Qin

doi:10.1145/3394171.3413579

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2009.02598 (eess)

[Submitted on 5 Sep 2020]

Title:Semi-supervised Multi-modal Emotion Recognition with Cross-Modal Distribution Matching

Authors:Jingjun Liang, Ruichen Li, Qin Jin

View PDF

Abstract:Automatic emotion recognition is an active research topic with wide range of applications. Due to the high manual annotation cost and inevitable label ambiguity, the development of emotion recognition dataset is limited in both scale and quality. Therefore, one of the key challenges is how to build effective models with limited data resource. Previous works have explored different approaches to tackle this challenge including data enhancement, transfer learning, and semi-supervised learning etc. However, the weakness of these existing approaches includes such as training instability, large performance loss during transfer, or marginal improvement.
In this work, we propose a novel semi-supervised multi-modal emotion recognition model based on cross-modality distribution matching, which leverages abundant unlabeled data to enhance the model training under the assumption that the inner emotional status is consistent at the utterance level across modalities.
We conduct extensive experiments to evaluate the proposed model on two benchmark datasets, IEMOCAP and MELD. The experiment results prove that the proposed semi-supervised learning model can effectively utilize unlabeled data and combine multi-modalities to boost the emotion recognition performance, which outperforms other state-of-the-art approaches under the same condition. The proposed model also achieves competitive capacity compared with existing approaches which take advantage of additional auxiliary information such as speaker and interaction context.

Comments:	10 pages, 5 figures, to be published on ACM Multimedia 2020
Subjects:	Audio and Speech Processing (eess.AS); Multimedia (cs.MM); Sound (cs.SD)
Cite as:	arXiv:2009.02598 [eess.AS]
	(or arXiv:2009.02598v1 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2009.02598
Related DOI:	https://doi.org/10.1145/3394171.3413579

Submission history

From: Ruichen Li [view email]
[v1] Sat, 5 Sep 2020 20:51:01 UTC (2,212 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Semi-supervised Multi-modal Emotion Recognition with Cross-Modal Distribution Matching

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Semi-supervised Multi-modal Emotion Recognition with Cross-Modal Distribution Matching

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators