acceptodds
Under review as a conference paper at ICLR 2027

Diagnosing and Mitigating Long-Session Degradation in Full-Duplex Speech Models

Abstract

End-to-end full-duplex spoken dialogue models enable natural, real-time spoken interactions with interruptions, backchannels, and overlapping speech, yet their reliability over extended sessions remains poorly understood. Existing evaluations have largely focused on short-horizon interactions, and even studies involving longer conversations have not systematically characterized how model behavior evolves over the course of a session. In this work, we address this gap by first characterizing how specific dialogue capabilities change over extended interactions and then developing methods to mitigate long-session degradation while improving the efficiency and generalization of extended-session inference. To characterize long-session behavior, we introduce a closed-loop evaluation framework for long-session stability and apply it to four publicly available full-duplex speech models in ten-minute conversations. Across models, we observe widespread degradation over time in responsiveness, content quality, and speech quality, highlighting the difficulty of maintaining stable conversational behavior over extended interactions. To mitigate this degradation, we propose a self-distillation-based training framework for lightweight LoRA adaptation that improves long-session stability with a small amount of additional training. We further introduce a two-resolution context (TRC) architecture that preserves recent interactions in full detail while compressing older interaction history into key-value representations, enabling models to learn how to effectively leverage both short-term and long-term context to maintain content and speech quality over longer sessions while reducing memory and computational costs. We apply our methods to three architectures with distinct formulations of full-duplex conversational context and evaluate long-session stability, information retention, and computational efficiency. Lightweight adaptation reduces performance degradation over ten-minute sessions across all three models. With TRC, the models also maintain these gains beyond the training horizon while reducing KV-cache growth by more than 90%

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.