MPE4G : Multimodal Pretrained Encoder for Co-Speech Gesture Generation

Gwantae Kim (Korea University); Seonghyeok Noh (Korea University); Insung Ham (Korea University); Hanseok Ko (Korea University)

DOI

SPS

Members: Free
IEEE Members: $11.00
Non-members: $15.00

07 Jun 2023

When virtual agents interact with humans, gestures are crucial to delivering their intentions with speech. Previous multimodal co-speech gesture generation models required encoded features of all modalities to generate gestures. If some input modalities are removed or contain noise, the model may not generate the gestures properly. To acquire robust and generalized encodings, we propose a novel framework with a multimodal pre-trained encoder for co-speech gesture generation. In the proposed method, the multi-head-attention-based encoder is trained with self-supervised learning to contain the information on each modality. Moreover, we collect full-body gestures that consist of 3D joint rotations to improve visualization and apply gestures to the extensible body model. Through the series of experiments and human evaluation, the proposed method renders realistic co-speech gestures not only when all input modalities are given but also when the input modalities are missing or noisy.

Tags:

Image and video synthesis, rendering, and visualization

MPE4G : Multimodal Pretrained Encoder for Co-Speech Gesture Generation

Gwantae Kim (Korea University); Seonghyeok Noh (Korea University); Insung Ham (Korea University); Hanseok Ko (Korea University)

Value-Added Bundle(s) Including this Product

IEEE ICASSP 2023, 4-10 June 2023, Greece. Virtual and In-Person Conference - Presentation Videos Product Bundle

More Like This

SVMV: SPATIOTEMPORAL VARIANCE-SUPERVISED MOTION VOLUME FOR VIDEO FRAME INTERPOLATION

Flow-Guided Deformable Alignment Network with Self-Supervision for Video Inpainting

Free-view Expressive Talking Head Video Editing

Join the IEEE Signal Processing Society