Full-duplex Spoken Language Modeling with Discrete Diffusion
Abstract
We introduce Diffuplex, a full-duplex spoken language model based on discrete diffusion that achieves an unprecedented balance between linguistic intelligence and response speed. With thinking enabled, Diffuplex finalizes an average of 10.15 tokens per forward pass, achieving substantially higher output speed than comparably intelligent turn-based autoregressive models. Without thinking, it matches the first-audio latency of existing full-duplex models while delivering markedly stronger spoken intelligence. This capability is enabled by separating when to say from what to say, rather than compressing both into a single autoregressive timeline. During inference, Diffuplex continuously tracks the conversation through an autoregressive interaction policy, while independently generating and refining response content through parallel diffusion. Consequently, semantic computation is no longer bounded by the perception audio frame rate, allowing intelligence and responsiveness to improve simultaneously. Taken together, these results establish a new paradigm for spoken agents that can sustain real-time dialogue while devoting parallel computation to intelligent response generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.