acceptodds
Under review as a conference paper at ICLR 2027

SpeechAnchor: Enhancing Speech Large Language Models for Consistent Multi-Turn Spoken Dialogue

Abstract

End-to-end Speech Large Language Models (SLMs) have shown strong potential for natural voice interaction by directly modeling speech-to-speech communication. However, practical voice interaction is inherently multi-turn, requiring SLMs to maintain semantic coherence, adapt to evolving paralinguistic cues, and preserve speaker style throughout the conversation. Existing SLMs remain limited in multi-turn dialogue because most training corpora are dominated by single-turn conversations, and existing architectures mainly rely on the LLM backbone to track information across dialogue history. To address these challenges, we propose **SpeechAnchor**, a multi-turn consistency enhancement framework for SLMs. SpeechAnchor combines a lightweight module that explicitly models cross-turn speech history for the speech decoder with a consistency-aware data pipeline that constructs semantically coherent, speaker-consistent, and paralinguistically rich multi-turn speech data. Experiments on three multi-turn and four single-turn speech benchmarks show consistent gains in multi-turn consistency without compromising general speech performance. These results establish SpeechAnchor as a practical framework for advancing multi-turn SLMs and provide a unified perspective to guide future research on multi-turn consistency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.