An Analysis Of Speech Enhancement And Recognition Losses In Limited Resources Multi-Talker Single Channel Audio-Visual Asr

Luca Pasa, Giovanni Morrone, Leonardo Badino

DOI

SPS

Members: Free
IEEE Members: $11.00
Non-members: $15.00

Length: 14:09

04 May 2020

In this paper, we analyzed how audio-visual speech enhancement can help to perform the ASR task in a cocktail party scenario. Therefore we considered two simple end-to-end LSTM-based models that perform single-channel audiovisual speech enhancement and phone recognition respectively. Then, we studied how the two models interact, and how to train them jointly affects the final result. We analyzed different training strategies that reveal some interesting and unexpected behaviors. The experiments show that during optimization of the ASR task the speech enhancement capability of the model significantly decreases and viceversa. Nevertheless the joint optimization of the two tasks shows a remarkable drop of the Phone Error Rate (PER) compared to the audio-visual baseline models trained only to perform phone recognition. We analyzed the behaviors of the proposed models by using two limited-size datasets, and in particular we used the mixed-speech versions of GRID and TCD-TIMIT.

Tags:

sps conference

icassp 2020 virtual conference

May 2020

icassp 2020

An Analysis Of Speech Enhancement And Recognition Losses In Limited Resources Multi-Talker Single Channel Audio-Visual Asr

Luca Pasa, Giovanni Morrone, Leonardo Badino

Value-Added Bundle(s) Including this Product

ICASSP 2020 Virtual Conference - Presentation Videos Product Bundle

More Like This

IEEE ICASSP 2023, 4-10 June 2023, Greece. Virtual and In-Person Conference - Presentation Videos Product Bundle

IEEE ICASSP 2024, 1 4-19 April 2024, Seoul, Korea. Conference Presentation Videos Bundle

ICIP 2022, October 16-19, 2022, Bordeaux, France - Presentation Videos Product Bundle

Join the IEEE Signal Processing Society