acceptodds
Under review as a conference paper at ICLR 2027

S-JEPA : Soft Clustering Anchors for Self-Supervised Speech Representation Learning

Abstract

Self-supervised speech encoders are predominantly trained by predicting discrete hard cluster IDs at masked positions, a recipe that collapses acoustic ambiguity at category boundaries and requires interrupting training to re-cluster the entire corpus between iterations. We introduce S-JEPA, a JEPA-style encoder-predictor pair trained to match the soft posteriors of a Gaussian Mixture Model at masked positions via KL divergence. Training runs as one continuous optimization trajectory in two phases: a fixed GMM over MFCC features, then an online GMM over encoder features, with the input layer selected adaptively from a label-free signal, so the re-cluster step runs online rather than offline and the choice of which transformer layer to cluster on is no longer hand-tuned. Under the SUPERB protocol, S-JEPA achieves the lowest WER among evaluated SSL methods below 90M parameters and matches HuBERT-Base on emotion recognition at roughly half its parameter count, with no offline re-clustering and no pretrained teacher. Pre-training data differs across the compared methods, so this is a comparison at matched parameter count rather than matched data or compute. An analysis of the predictor’s per-frame entropy on held-out speech reveals a bimodal distribution with a substantial minority of frames near the entropy of a perfect two-cluster tie, showing that the objective retains inter-cluster mass on frames where a hard target would have to discard it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.