acceptodds
Under review as a conference paper at ICLR 2027

DeIRD: A Decoupled Continuous Implicit Reasoning Framework for Dialogue Models

Abstract

By leveraging the ”think before speaking” response paradigm, large models can generate a Chain-of-Thought (CoT) prior to answering to enhance their deep reasoning capabilities. However, directly applying this paradigm to spoken dialogues poses three severe challenges: huge training costs caused by the modality gap, the unacceptable latency introduced by the ”think before speaking” mechanism, and limited reasoning exploration depth. To address these issues, we propose DeIRD, a latency-controllable and low-training-cost implicit reasoning framework for spoken dialogue models that simultaneously achieves ”think and speak decoupled” and ”think before speaking”. This framework is primarily realized through three core components: 1) We train a text VAE to compress variable-length textual CoT sequences into fixed-length continuous latents, serving as the carrier for implicit reasoning; 2) We introduce a continuous masked prediction model DMIR that takes user speech as input and outputs a fixed-length sequence of these latent targets via iterative confidence-guided ”mask-predict” decoding, to achieve ”think and speak decoupled” and ”think before speaking”; 3) We construct an implicit reasoning training dataset, LatentReason-40k, and find that the downstream backbones just require one epoch of end-to-end fine-tuning to acquire superior reasoning capabilities compared to explicit reasoning models trained on the same volume of data. Compared with various explicit-CoT baselines, DeIRD notably reduces the Time to First Audio (TTFA) while improving general dialogue quality, style control, deep reasoning, and reasoning diversity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.