acceptodds
Under review as a conference paper at ICLR 2027

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding

Abstract

Neural audio codecs are a key component of speech processing pipelines, compressing audio into discrete tokens for downstream modeling. However, existing codecs struggle to balance reconstruction quality with token efficiency, often encoding perceptually irrelevant information such as background noise and recording artifacts at the expense of linguistically and acoustically meaningful content. We reframe audio tokenization as a selective information bottleneck problem and propose CleanCodec, a neural audio codec utilizing a denoising training objective, where the codec receives stochastically degraded speech as input but is trained to reconstruct the corresponding clean signal. CleanCodec learns to encode only perceptually important features and achieves state-of-the-art tokenization efficiency at just 12.5 t/s, improving speaker similarity by 0.21 and reducing word error rate by up to 16.3% compared to existing codecs. Evaluations on downstream text-to-speech and voice conversion tasks further demonstrate improved performance and up to 17x faster inference, highlighting significant efficiency gains due to reduced token rate.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.