Bringing Characters into Conversation: Natural and Emotion-Aware Spoken Role-Play
Abstract
Role-play models can reproduce a character's biography yet struggle to sustain a natural conversation. Spoken interaction also requires responses that account for the user's vocal cues and express empathy in a way that fits the persona. We adapt Fun-Audio-Chat-8B through continued supervised fine-tuning (cSFT) on persona-conditioned replies paired with real and synthetic user speech, followed by group relative policy optimization (GRPO). A history-aware reward rubric scores persona fidelity, contextual relevance, semantic and structural repetition, and speech-cue use. Evaluation covers text-to-text (T2T), speech-to-text (S2T), and speech-to-speech (S2S) responses. Expression naturalness (EN) improves by 0.89–1.29 points over the base model across three T2T evaluations and a supplementary S2T test. On an in-house S2S test of 100 sessions, each with 20–50 exchanges about fictional characters' experiences, EN rises from 1.02 to 3.02. GRPO raises mean empathy scores by 0.08–0.52 points over cSFT across four S2T datasets. On the complete VStyle Implicit Empathy test, which evaluates content and vocal delivery jointly, the score increases from 3.73 to 4.33. Human EN means follow the same checkpoint ordering as the model judge. Reward ablations reveal trade-offs between persona consistency and empathy rather than uniform gains from every criterion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.