acceptodds
Under review as a conference paper at ICLR 2027

CueDuet: Distilling What to Say and How to Speak for Paralinguistic-Aware Spoken Dialogue

Abstract

Speech language models have made substantial progress in spoken interaction, yet adapting response content and vocal delivery to paralinguistic cues remains challenging. Teacher-guided adaptation can transfer cue-informed dialogue behavior, but its supervision mixes general response capabilities with the local adjustments needed for paralinguistic cues. Moreover, appropriate wording alone does not ensure that the generated speech conveys a suitable vocal style. We propose CueDuet, a dual-stream on-policy distillation framework that addresses these complementary needs through teacher supervision on student-generated responses. For the text stream, a single teacher evaluates the same response prefix from cue-aware and cue-agnostic perspectives; divergence between its predictions determines where to emphasize supervision for paralinguistic adaptation. For the speech stream, a speech teacher adapted on expressive text–speech pairs conditions on the student's reply, guiding speech generation toward faithful content and an appropriate style. The evaluation covers emotion- and age-sensitive dialogue across three benchmarks and two architecturally distinct speech language models, comparing offline distillation and standard on-policy distillation baselines. Across both backbones, CueDuet improves cue-responsive content and vocal delivery over offline and standard on-policy distillation, with particularly strong gains in cue-grounded safety. Text-loss analyses show that CueDuet consistently concentrates supervision on cue-responsive positions throughout training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.