VoiceDiff: Explaining Voice Differences with Speech Language Models
Abstract
As real-time spoken dialogue and multi-party assistants expand voice-based interaction, speech language models need to understand information about speakers beyond verbal content and use it in their responses. Because one person can sound different across utterances and different people can share similar characteristics, voice similarity must be distinguished from speaker identity. We introduce Voice Difference Explanation, a task covering speaker profiles such as gender and age, voice characteristics such as pitch and timbral brightness, and recording conditions such as noise and reverberation. It connects individual voice judgments to comparisons of shared and differing attributes, integrating them into a natural-language response alongside a separate identity verdict. Public speech metadata, acoustic measurements, and controlled recording conditions provide training targets and VoiceDiff-Bench. To study useful representations for these judgments, we introduce VoiceDiff, a reference model whose hierarchical speaker tokenizer combines a frozen speaker encoder's pooled speaker embedding with temporally resolved frame-level features. Task-specific supervision improves general Audio-LLMs on multiple voice judgments, while VoiceDiff achieves stronger pitch, timbral-brightness, and identity results than the evaluated task-adapted general models. At a matched audio-token budget, speaker embeddings favor identity judgments, frame-level features favor detailed acoustic judgments, and their combination supports both. Relation-description training integrates comparisons, summaries, and identity into one response with category-summary accuracy generally close to rule-based composition of individual outputs. Overall conclusions largely follow the comparisons stated in the response; error analysis identifies constituent comparison accuracy as a key basis for more accurate relation descriptions. The benefits of task-specific supervision also extend to selected external cases, unseen during task training, in which spoken claims conflict with vocal cues. These findings support task-specific supervision and complementary speech representations for grounding linguistic judgments and explanations in the voice itself.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.