Libri-Adapt: A New Speech Dataset For Unsupervised Domain Adaptation

Akhil Mathur, Fahim Kawsar, Nadia Berthouze, Nicholas Lane

DOI

SPS

Members: Free
IEEE Members: $11.00
Non-members: $15.00

Length: 13:47

04 May 2020

This paper introduces a new dataset, Libri-Adapt, to support unsupervised domain adaptation research on speech recognition models. Built on top of the LibriSpeech corpus, Libri-Adapt contains 7200 hours of English speech recorded on mobile and embedded-scale microphones, and spans 72 different domains that are representative of the challenging practical scenarios encountered by ASR models. More specifically, Libri-Adapt facilitates the study of domain shifts in ASR models caused by a) different acoustic environments, b) variations in speaker accents, c) previously unexplored factors such as heterogeneity in the hardware and platform software of the microphones, and d) a combination of the aforementioned three shifts. We also provide a number of baseline results quantifying the impact of these domain shifts on the Mozilla DeepSpeech2 ASR model.

Tags:

sps conference

icassp 2020 virtual conference

May 2020

icassp 2020

Libri-Adapt: A New Speech Dataset For Unsupervised Domain Adaptation

Akhil Mathur, Fahim Kawsar, Nadia Berthouze, Nicholas Lane

Value-Added Bundle(s) Including this Product

ICASSP 2020 Virtual Conference - Presentation Videos Product Bundle

More Like This

IEEE ICASSP 2023, 4-10 June 2023, Greece. Virtual and In-Person Conference - Presentation Videos Product Bundle

IEEE ICASSP 2024, 1 4-19 April 2024, Seoul, Korea. Conference Presentation Videos Bundle

ICIP 2022, October 16-19, 2022, Bordeaux, France - Presentation Videos Product Bundle

Join the IEEE Signal Processing Society