acceptodds
Under review as a conference paper at ICLR 2027

Laconic: Configurable low-rate speech representations

Abstract

Speech representations spend capacity along two axes: how many tokens are emitted per second, and how many dimensions each token carries. Yet they are usually characterized by token rate alone, with token dimensionality treated as a fixed architectural choice. We ask how capacity should be allocated between the two. We introduce Laconic, a continuous speech representation whose token rate is chosen at encoding time and whose token dimensionality is chosen at reading time, as a nested prefix of each token, so that a single encoder exposes a two-axis family of operating points. Mapping this space, we find that token rate and dimensionality are partially interchangeable between roughly 3 and 6 tokens per second, but the exchange reaches a temporal floor at 1 token per second, where no tested dimensionality recovers the lost content. Different information also occupies different regions of the space. Under our nested objective, linguistic content becomes readable from low-dimensional prefixes while speaker identity emerges only at high-dimensional ones, and is carried more compactly by a small set of utterance-level tokens. The same geometry governs generation. In zero-shot TTS, low-dimensional tokens are easiest to generate at higher token rates, high-dimensional tokens help as the token rate decreases, and speaker similarity and perceptual quality favor different dimensionalities. These results suggest treating speech representation capacity not as a single rate budget, but as an allocation across token rate and dimensionality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.