acceptodds
Under review as a conference paper at ICLR 2027

OLIVE: View-Augmented Latent Prediction with Waveform Reconstruction for Speech SSL

Abstract

We propose Online Latent prediction with Invariant Views and rEconstruction (OLIVE), a self-supervised speech representation learning framework that jointly optimizes analysis and synthesis objectives. OLIVE couples masked latent prediction and waveform reconstruction through a functional separation within a single encoder: reconstruction constrains early encoder features to retain signal-level information, while masked latent prediction shapes later contextual representations for robust downstream performance, with waveform view augmentations defining the invariances they learn. To the best of our knowledge, OLIVE is the first speech SSL model that directly reconstructs waveforms, retains its vocoder after pre-training, and remains competitive on discriminative tasks. Adding reconstruction largely preserves the speech recognition performance of masked latent prediction alone, while improving waveform reconstruction and speaker tasks. Adding waveform mixup and gain augmentations provides a controllable way to shift the representations further toward speaker, generation, and reconstruction tasks, while remaining competitive on recognition and semantic tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.