(Slides) Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition

Qiuqiang Kong

SPS

Members: Free
IEEE Members: $11.00
Non-members: $15.00

Pages/Slides: 49

29 Feb 2024

Audio pattern recognition is an important research topic in the machine learning area, and includes several tasks such as audio tagging, acoustic scene classification, music classification, speech emotion classification and sound event detection. In this blog, we introduce pretrained audio neural networks (PANNs) trained on the large-scale AudioSet dataset. These PANNs are transferred to other audio related tasks. We investigate the performance and computational complexity of PANNs modeled by a variety of convolutional neural networks. We propose an architecture called Wavegram-Logmel-CNN using both log-mel spectrogram and waveform as input feature. Our best PANN system achieves a state-of-the-art mean average precision (mAP) of 0.439 on AudioSet tagging, outperforming the best previous system of 0.392. We transfer PANNs to six audio pattern recognition tasks, and demonstrate state-of-the-art performance in several of those tasks.

Tags:

IEEE sps webinars

webinars

2024 webinars

audio tagging

pretrained audio neural networks

transfer learning

(Slides) Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition

Qiuqiang Kong

More Like This

Short Course Bundle: ICIP 2023 COURSE 2: Short Course: Unboxing Advancements in Biomedical Image Processing (Parts 1-4)

(Slides) Joint Waveform and Beamforming Design for RIS-Aided ISAC Systems

Joint Waveform and Beamforming Design for RIS-Aided ISAC Systems

Join the IEEE Signal Processing Society