Told, Then Aware: On-Policy Distillation for Empathetic Speech-to-Speech Dialogue
Abstract
The same sentence spoken in a different tone calls for a different reply, yet end-to-end speech-to-speech (S2S) models usually respond to the words and largely ignore the voice, even when the delivery contradicts what is said. The common remedy is more empathetic dialogue data with paired annotations, but that data is scarce. We find the solution already latent in the model: a pretrained S2S model can carry emotion in its own speech, just not always the emotion empathy calls for, and it can answer empathetically both in content and in how it sounds as soon as the user's affect is explicitly stated in words. What is missing is not the capability but an objective that requires the voice to change the reply. We therefore adopt on-policy distillation (OPD): the teacher is prompted with the user's affect information as a training-time privilege, and this frozen privileged teacher provides dense per-position distributional supervision over both text and speech positions along the student's own sampled response trajectories. On SpeechParaling-Bench, our method on one backbone improves the Chinese score from 43.6 to 56.3 and the English score from 60.5 to 71.8, outperforming all evaluated systems. We observe consistent gains on EchoMind, VocalBench-zh, and with an alternative backbone. Furthermore, human evaluation confirms that the model trained via our approach achieves the best performance, demonstrating OPD’s potential for generating empathetic spoken responses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.