A Comparison of Semi-Supervised Learning Techniques for Streaming ASR at Scale

Charles C Peyser (Google Inc.); Michael Picheny (NYU); Kyunghyun Cho (New York University); Tara Sainath (Google); W. Ronny Huang (Google); Rohit Prabhavalkar (Google)

DOI

SPS

Members: Free
IEEE Members: $11.00
Non-members: $15.00

06 Jun 2023

Unpaired text and audio injection have emerged as dominant methods for improving ASR performance in the absence of a large labeled corpus. However, little guidance exists on deploying these methods to improve production ASR systems that are trained on very large supervised corpora and with realistic requirements like a constrained model size and CPU budget, streaming capability, and a rich lattice for rescoring and for downstream NLU tasks. In this work, we compare three state-of-the-art semi-supervised methods encompassing both unpaired text and audio as well as several of their combinations in a controlled setting using joint training. We find that in our setting these methods offer many improvements beyond raw WER, including substantial gains in tail-word WER, decoder computation during inference, and lattice density.

Tags:

Large vocabulary continuous speech recognition/search

A Comparison of Semi-Supervised Learning Techniques for Streaming ASR at Scale

Charles C Peyser (Google Inc.); Michael Picheny (NYU); Kyunghyun Cho (New York University); Tara Sainath (Google); W. Ronny Huang (Google); Rohit Prabhavalkar (Google)

Value-Added Bundle(s) Including this Product

IEEE ICASSP 2023, 4-10 June 2023, Greece. Virtual and In-Person Conference - Presentation Videos Product Bundle

More Like This

Effectiveness of Mining Audio and Text Pairs from Public Data for Improving ASR Systems for Low-Resource Languages

ROBUST ACOUSTIC AND SEMANTIC CONTEXTUAL BIASING IN NEURAL TRANSDUCERS FOR SPEECH RECOGNITION

Large-scale Language Model Rescoring on Long-form Data

Join the IEEE Signal Processing Society