Self-supervised learning of audio representations using angular contrastive loss

Shanshan Wang (Tampere University); Soumya Tripathy (Tampere University of Technology); Annamaria Mesaros (Tampere University)

DOI

SPS

Members: Free
IEEE Members: $11.00
Non-members: $15.00

07 Jun 2023

In Self-Supervised Learning (SSL), various pretext tasks are designed for learning feature representations through contrastive loss. However, previous studies have shown that this loss is less tolerant to semantically similar samples due to the inherent defect of instance discrimination objectives, which may harm the quality of learned feature embeddings used in downstream tasks. To improve the discriminative ability of feature embeddings in SSL, we propose a new loss function called Angular Contrastive Loss (ACL), a linear combination of angular margin and contrastive loss. ACL improves contrastive learning by explicitly adding an angular margin between positive and negative augmented pairs in SSL. Experimental results show that using ACL for both supervised and unsupervised learning significantly improves performance. We validated our new loss function using the FSDnoisy18k dataset, where we achieved 73.6% and 77.1% accuracy in sound event classification using supervised and self-supervised learning, respectively.

Tags:

Modeling, analysis and synthesis of acoustic environments

Self-supervised learning of audio representations using angular contrastive loss

Shanshan Wang (Tampere University); Soumya Tripathy (Tampere University of Technology); Annamaria Mesaros (Tampere University)

Value-Added Bundle(s) Including this Product

IEEE ICASSP 2023, 4-10 June 2023, Greece. Virtual and In-Person Conference - Presentation Videos Product Bundle

More Like This

Neural Fourier Shift for Binaural Speech Rendering

Lightweight Annotation and Class Weight Training for Automatic Estimation of Alarm Audibility in Noise

SUBBAND DEPENDENCY MODELING FOR SOUND EVENT DETECTION

Join the IEEE Signal Processing Society