AN ASR-FREE FLUENCY SCORING APPROACH WITH SELF-SUPERVISED LEARNING

Wei Liu (The Chinese University of Hong Kong); Kaiqi Fu (Bytedance); Xiaohai Tian (ByteDance); Shuju Shi (ByteDance); Wei Li (Bytedance); Zejun Ma (Bytedance); Tan Lee (The Chinese University of Hong Kong)

DOI

SPS

Members: Free
IEEE Members: $11.00
Non-members: $15.00

06 Jun 2023

A typical fluency scoring system generally relies on an automatic speech recognition (ASR) system to obtain time stamps in input speech for subsequent calculation of fluency related features or directly modeling speech fluency with an end-to-end approach. This paper describes a novel ASR-free approach for automatic fluency assessment using self-supervised learning (SSL). Specifically, wav2vec2.0 is used to extract frame-level speech features, followed by K-means clustering to assign a pseudo label (cluster index) to each frame. A BLSTM-based model is trained to predict an utterance-level fluency score from frame-level SSL features and the corresponding cluster indexes. Neither speech transcription nor time stamp information is required in the proposed system. It is ASR-free and can potentially avoid the ASR errors effect in practice. Experimental results carried out on non-native English databases show that the proposed approach significantly improves the performance in the "open response" scenario as compared to previous methods and matches the recently reported performance in the "read aloud" scenario.

Tags:

Machine learning methods for language

AN ASR-FREE FLUENCY SCORING APPROACH WITH SELF-SUPERVISED LEARNING

Wei Liu (The Chinese University of Hong Kong); Kaiqi Fu (Bytedance); Xiaohai Tian (ByteDance); Shuju Shi (ByteDance); Wei Li (Bytedance); Zejun Ma (Bytedance); Tan Lee (The Chinese University of Hong Kong)

Value-Added Bundle(s) Including this Product

IEEE ICASSP 2023, 4-10 June 2023, Greece. Virtual and In-Person Conference - Presentation Videos Product Bundle

More Like This

UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction

A Sentiment and Syntactic-Aware Graph Convolutional Network for Aspect-level Sentiment Classification

SELF SUPERVISED BERT FOR LEGAL TEXT CLASSIFICATION

Join the IEEE Signal Processing Society