REPEAT AFTER ME: SELF-SUPERVISED LEARNING OF ACOUSTIC-TO-ARTICULATORY MAPPING BY VOCAL IMITATION

Marc-Antoine Georges, Laurent Girin, Jean-Luc Schwartz, Thomas Hueber, Julien Diard

DOI

SPS

Members: Free
IEEE Members: $11.00
Non-members: $15.00

Length: 00:08:50

13 May 2022

We propose a computational model of speech production combining a pre-trained neural articulatory synthesizer able to reproduce complex speech stimuli from a limited set of interpretable articulatory parameters, a DNN-based internal forward model predicting the sensory consequences of articulatory commands, and an internal inverse model based on a recurrent neural network recovering articulatory commands from the acoustic speech input. Both forward and inverse models are jointly trained in a self-supervised way from raw acoustic-only speech data from different speakers. The imitation simulations are evaluated objectively and subjectively and display quite encouraging performances.

Tags:

computational models

articulatory synthesis

speech production

representation learning

REPEAT AFTER ME: SELF-SUPERVISED LEARNING OF ACOUSTIC-TO-ARTICULATORY MAPPING BY VOCAL IMITATION

Marc-Antoine Georges, Laurent Girin, Jean-Luc Schwartz, Thomas Hueber, Julien Diard

Value-Added Bundle(s) Including this Product

ICASSP 2022, May 2022 Virtual and In-Person Conference - Presentation Videos Product Bundle

More Like This

Tutorial: Understanding Deep Representation Learning via Neural Collapse

Slides: The Changing Landscape of Speech Foundation Models

The Changing Landscape of Speech Foundation Models

Join the IEEE Signal Processing Society