Interlude: Reasoning Between Tokens in Vision-Language Models
Abstract
Multimodal reasoning mixes two kinds of intermediate steps: symbolic relations that language expresses precisely and continuous visual judgments that text can only approximate. Approaches using a single medium face a trade-off: explicit chain-of-thought can discard fine visual details, while purely latent reasoning can lose the readable scaffold. We propose **Interlude**, retaining language as the backbone while punctuating it with latent segments: short, undecoded spans of recurrent hidden-state computation used by subsequent text. We seek states that remain useful to subsequent text while retaining visual evidence that is difficult to verbalize. Motivated by information under-saturation and the modality gap, we combine semantic guidance on the language-head readout with a geometric margin favoring an intermediate reference over either modal pole. Interlude improves over Qwen3-VL-8B-Instruct on all six evaluated benchmarks, with an average gain of 4.72%, and achieves the highest Overall score among the compared explicit and latent reasoning baselines. Its Overall score also exceeds that of Qwen3-VL-8B-Thinking while using just 7.6% of its decoded tokens in our efficiency evaluation. Geometric analysis shows a pronounced increase in visual alignment for the component orthogonal to the word-embedding principal subspace, with little change within that subspace.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.