acceptodds
Under review as a conference paper at ICLR 2027

What Did I Actually Say? Self-Listening for Full-Duplex Speech Models

Abstract

Full-duplex spoken language models can listen and speak simultaneously, enabling them to handle interruptions and backchannels in human conversation. However, text generation, speech synthesis, and audio playback proceed asynchronously. As a result, what a model believes it has said may not match what has actually been played to the user. We refer to the problem of recovering from an interruption while remaining aware of the assistant's realized speech as anchor interruption. To address this problem, we propose self-listening, a full-duplex modeling approach that interleaves user speech, assistant text, and the assistant's played speech. By feeding the realized speech output back to the model as an input stream, self-listening grounds interruption recovery in what the user has actually heard. We further introduce AnchorSpeech, which contains both training data and a controlled evaluation set for fine-grained anchoring. AnchorSpeech evaluates whether a model can accurately identify the spoken boundary reached before an interruption. Experiments show that, compared with a full-duplex baseline, models equipped with self-listening achieve better anchoring performance and produce more coherent responses after interruption.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.