What Makes a Good Full-Duplex SLM? A Systematic Study of Architecture, Objectives, and Data
Abstract
Full-Duplex Speech Language Models (FD-SLMs) model user and assistant speech as separate channels, allowing a spoken assistant to listen while speaking and handle overlapping speech. However, existing systems combine different architectures, backbones, objectives, and data recipes, which makes their design choices difficult to compare. We conduct controlled experiments on four axes—dual-channel sequence layout and speech granularity, duplex training objectives, text guidance granularity, and real versus synthetic duplex data—using a shared pipeline and two open SLM backbones. Interleaved training improves mean dialogue quality and interruption recovery over our step-aligned implementation across both backbones and all tested speech chunk sizes. Joint user–assistant prediction with equal loss weights increases assistant-speech cross-entropy and usually reduces post-interruption latency, but its effects on dialogue quality and post-interruption recovery rate depend on the backbone and chunk size. Turn-level text guidance favors mean dialogue quality and faster post-interruption responses, whereas chunk-level guidance makes models easier to interrupt. Synthetic dialogue improves spoken QA on GLM-4-Voice and post-interruption latency and response rating on both backbones, but reduces post-interruption recovery rate; we also observe text-channel collapse in the synthetic-trained Kimi-Audio checkpoint. These findings provide actionable insights into selecting architectures, training objectives, and dialogue data to balance response quality, generation stability, and turn-taking behavior.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.