acceptodds
Under review as a conference paper at ICLR 2027

Where codecs draw the line: allocation profiles for speech tokenisers

Abstract

Quantising a speech representation is dividing it: a codebook partitions latent space, and every distinction finer than a cell is destroyed. Analyses of speech tokenisers ask what their codes retain, not where their boundaries fall. We charac- terise a partition by the reconstruction error each phonetic category incurs — an allocation profile, comparable where the partitions themselves are not. It separates two things a single test conflates: added levels may narrow the gaps between cate- gories, or change which ones are finely resolved. Across four widely used codecs and one we train to reverse a single architectural choice, depth does the first, and only quantisation in a learnt low-dimensional projection does the second. And what the first level resolves finely is largely a restatement of the signal: energy, duration and frequency account for 82.2% of one codec’s acoustic profile against 16.2% of the distilled stream trained beside it, which allocates along unrelated lines. Depth mostly refines an allocation that architecture has already settled.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.