acceptodds
Under review as a conference paper at ICLR 2027

SLIMDUCER: A STREAMABLE SPEECH LLM WITH 3M TRAINABLE PARAMETERS

Abstract

Speech is normally given to a language model at its input, where encoder features occupy context positions the model was never pre-trained to read. Slimducer attaches at the other end: a 3 M-parameter joint network writes into the language model’s own frozen output projection, so the model sees only its own vocabulary indices, beginning at bos, as in pre-training. Nothing else trains — 0.38% of 783.7 M, with the speech encoder and the language model frozen and its 151,936- way head unchanged — and every model reported here was trained on a single 24GB consumer GPU. Keeping that head whole is what removing the label axis from the alignment lattice buys: no T × L × |V| tensor is retained for a backward pass, where predictor-side transducers must cut their vocabulary to a few thousand. Decoding is framesynchronous, so bounded audio chunks reproduce whole-utterance decoding — streaming that an input-side coupling cannot offer at all, having to see the utterance before the language model attends over it. The same 3M parameters carry beyond transcription. On LibriSpeech Slimducer is slightly behind a fully fine-tuned wav2vec 2.0 of thirty times its trainable size and 18% to 39% better relative outside that domain; with a 3.4M text-only adapter on the frozen backbone it translates speech into German at 14.98 BLEU against 12.09 for a cascade of the same backbone, and answers questions about the utterance just recognised.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.