Quantifying Separation in Representations for Speech Tasks
Abstract
Self-supervised speech modelling aims to learn useful general-purpose speech representations for a wide variety of downstream tasks. Despite their superior capability to capture complex data properties, one of the drawbacks is that information relevant for different tasks is typically entangled in the representation space, thereby limiting general task performance and imposing extra difficulty for universal speech representation learning. A formal measure of the degree of separation of task-relevant information in representations has the potential to mitigate this but is under-investigated in the literature. To address this gap, this work introduces the concept of Task Separation, which extends the concept of disentanglement from latent factors to downstream task use. It aims to quantify the degree of separation of representation content between tasks. Using Self-supervised Learning (SSL) features linear models are added to probing tasks. An information-theoretic measure of Task Separation is derived from the parameter space of such linear probers. Experiments examine the degree of separation between four speech tasks and different SSL presentations: speaker identification (SID), speech emotion recognition (SER), phone recognition (PR) and automatic speech recognition (ASR). Experiments show that: 1) the separation score for tasks between SID and SER, PR and ASR, respectively, is at a moderate level, ranging from 0.32 to 0.67, suggesting the difficulty of separating task-relevant information of SID from the other three tasks; 2) with a relatively high separation score, the linear model parameters from SER, PR and ASR can filter out the relevant information for SID effectively; 3) the distribution derived from the parameter space of linear probers can be used to interpret the optimal design of a minimum hierarchical model to combine two speech tasks. These findings enrich the scientific understanding of speech representations from the view of downstream tasks, with a formal evaluation protocol contributing to the goal of universal speech representation learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.