acceptodds
Under review as a conference paper at ICLR 2027

Bringing Characters into Conversation: Natural and Emotion-Aware Spoken Role-Play

Abstract

Role-play models can reproduce a character's biography yet struggle to sustain a natural conversation. Spoken interaction also requires responses that account for the user's vocal cues and express empathy in a way that fits the persona. We adapt Fun-Audio-Chat-8B through continued supervised fine-tuning (cSFT) on persona-conditioned replies paired with real and synthetic user speech, followed by group relative policy optimization (GRPO). A history-aware reward rubric scores persona fidelity, contextual relevance, semantic and structural repetition, and speech-cue use. Evaluation covers text-to-text (T2T), speech-to-text (S2T), and speech-to-speech (S2S) responses. Expression naturalness (EN) improves by 0.89–1.29 points over the base model across three T2T evaluations and a supplementary S2T test. On an in-house S2S test of 100 sessions, each with 20–50 exchanges about fictional characters' experiences, EN rises from 1.02 to 3.02. GRPO raises mean empathy scores by 0.08–0.52 points over cSFT across four S2T datasets. On the complete VStyle Implicit Empathy test, which evaluates content and vocal delivery jointly, the score increases from 3.73 to 4.33. Human EN means follow the same checkpoint ordering as the model judge. Reward ablations reveal trade-offs between persona consistency and empathy rather than uniform gains from every criterion.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.