acceptodds
Under review as a conference paper at ICLR 2027

Eliciting Latent Capabilities in Spoken Dialogue Models via On-Policy Self-Distillation

Abstract

Post-training has become increasingly important for unlocking and refining the capabilities of foundation models. Spoken dialogue models introduce an additional challenge, as post-training should improve semantic reasoning and paralinguistic responsiveness jointly. We find that both are limited not by a lack of capability but by insufficient elicitation from speech alone. Providing transcripts and structured rationales substantially recovers both. Further analysis reveals local redundancy in discrete audio representations. Based on these findings, we propose VoxOPD, a cross-modal hierarchical on-policy self-distillation framework for spoken dialogue models. During training, the student generates response trajectories from speech alone, while a frozen teacher initialized from the same model conditions on privileged context to provide dense supervision along those trajectories. Beyond token-level distribution alignment, the framework assigns shared credit to audio segments through chunk-level supervision, accommodating differences in information density and local redundancy between text and audio representations. This hierarchical supervision transfers the capabilities elicited under augmented contexts to the speech-only setting. On Step-Audio 2 mini, VoxOPD improves the VoiceBench capability-specific task average by 4.0 and VStyle by 0.71. Further experiments on Kimi-Audio and VITA-Audio consistently demonstrate its effectiveness in improving both reasoning and spoken interaction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.