MM-DFN: Multimodal Dynamic Fusion Network For Emotion Recognition in Conversations

Dou Hu, Xiaolong Hou, Lianxin Jiang, Lingwei Wei, Yang Mo

DOI

SPS

Members: Free
IEEE Members: $11.00
Non-members: $15.00

Length: 00:09:56

10 May 2022

Emotion Recognition in Conversations (ERC) has considerable prospects for developing empathetic machines. For multimodal ERC, it is vital to understand context and fuse modality information in conversations. Recent graph-based fusion methods generally aggregate multimodal information by exploring unimodal and cross-modal interactions in a graph. However, they accumulate redundant information at each layer, limiting the context understanding between modalities. In this paper, we propose a novel Multimodal Dynamic Fusion Network (MM-DFN) to recognize emotions by fully understanding multimodal conversational context. Specifically, we design a new graph-based dynamic fusion module to fuse multimodal context features in a conversation. The module reduces redundancy and enhances complementarity between modalities by capturing the dynamics of contextual information in different semantic spaces. Extensive experiments on two public benchmark datasets demonstrate the effectiveness and superiority of the proposed model.

Tags:

emotion recognition in conversations

multimodal fusion

emotion recognition

dialogue systems

MM-DFN: Multimodal Dynamic Fusion Network For Emotion Recognition in Conversations

Dou Hu, Xiaolong Hou, Lianxin Jiang, Lingwei Wei, Yang Mo

Value-Added Bundle(s) Including this Product

ICASSP 2022, May 2022 Virtual and In-Person Conference - Presentation Videos Product Bundle

More Like This

MODALITY-AWARE OOD SUPPRESSION USING FEATURE DISCREPANCY FOR MULTI-MODAL EMOTION RECOGNITION

Audio-Visual Quality Assessment for User Generated Content: Database and Method

CUSTOMER SATISFACTION ESTIMATION USING UNSUPERVISED REPRESENTATION LEARNING WITH MULTI-FORMAT PREDICTION LOSS

Join the IEEE Signal Processing Society