Internal Language Model Estimation based Adaptive Language Model Fusion for Domain Adaptation

Rao Ma (University of Cambridge); Xiaobo Wu (ByteDance); Jin Qiu (ByteDance); Yanan Qin (ByteDance); Haihua Xu (ByteDance); Peihao Wu (Bytedance); Zejun Ma (Bytedance)

DOI

SPS

Members: Free
IEEE Members: $11.00
Non-members: $15.00

07 Jun 2023

ASR model deployment environment is ever-changing, and the incoming speech can be switched across different domains during a session. This brings a challenge for effective domain adaptation when only target domain text data is available, and our objective is to obtain obviously improved performance on the target domain while the performance on the general domain is less undermined. In this paper, we propose an adaptive LM fusion approach called internal language model estimation based adaptive domain adaptation (ILME-ADA). To realize such an ILME-ADA, an interpolated log-likelihood score is calculated based on the maximum of the scores from the internal LM and the external LM (ELM) respectively. We demonstrate the efficacy of the proposed ILME-ADA method with both RNN-T and LAS modeling frameworks employing neural network and n-gram LMs as ELMs respectively on two domain specific (target) test sets. The proposed method can achieve significantly better performance on the target test sets while it gets minimal performance degradation on the general test set, compared with both shallow and ILME-based LM fusion methods.

Tags:

language modeling

Internal Language Model Estimation based Adaptive Language Model Fusion for Domain Adaptation

Rao Ma (University of Cambridge); Xiaobo Wu (ByteDance); Jin Qiu (ByteDance); Yanan Qin (ByteDance); Haihua Xu (ByteDance); Peihao Wu (Bytedance); Zejun Ma (Bytedance)

Value-Added Bundle(s) Including this Product

IEEE ICASSP 2023, 4-10 June 2023, Greece. Virtual and In-Person Conference - Presentation Videos Product Bundle

More Like This

Large-Scale and Parameter-Efficient Language Modeling for Speech Processing

HAG: Hierarchical Attention with Graph Network for Dialogue Act Classification in Conversation

Enhancing Unsupervised Speech Recognition with Diffusion GANs

Join the IEEE Signal Processing Society