Thinking in the Gaps: Reusing the Text–Speech Frame-Rate Gap for Chain-of-Thought Reasoning
Abstract
Spoken language models (SLMs) enable natural human-computer interaction, but most decode text and speech at different frame rates on a shared autoregressive grid, leaving many text positions idle while speech generation continues. We introduce Thinking in the Gaps, which reuses these idle positions for stepwise reasoning in parallel with speech generation. Because full reasoning traces often exceed the available gap capacity, a context-conditioned compressor constructs length-constrained reasoning targets that fit within these gaps. Across five spoken mathematical question-answering benchmarks, this method improves average accuracy over the matched baseline at every tested frame rate, with the largest gain of 14.05 percentage points at 25 Hz. Crucially, these gains come at virtually no additional inference cost and without any interruption to speech generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.