SpeechAnchor: Enhancing Speech Large Language Models for Consistent Multi-Turn Spoken Dialogue
Abstract
End-to-end Speech Large Language Models (SLMs) have shown strong potential for natural voice interaction by directly modeling speech-to-speech communication. However, practical voice interaction is inherently multi-turn, requiring SLMs to maintain semantic coherence, adapt to evolving paralinguistic cues, and preserve speaker style throughout the conversation. Existing SLMs remain limited in multi-turn dialogue because most training corpora are dominated by single-turn conversations, and existing architectures mainly rely on the LLM backbone to track information across dialogue history. To address these challenges, we propose **SpeechAnchor**, a multi-turn consistency enhancement framework for SLMs. SpeechAnchor combines a lightweight module that explicitly models cross-turn speech history for the speech decoder with a consistency-aware data pipeline that constructs semantically coherent, speaker-consistent, and paralinguistically rich multi-turn speech data. Experiments on three multi-turn and four single-turn speech benchmarks show consistent gains in multi-turn consistency without compromising general speech performance. These results establish SpeechAnchor as a practical framework for advancing multi-turn SLMs and provide a unified perspective to guide future research on multi-turn consistency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.