acceptodds
Under review as a conference paper at ICLR 2027

DuplexKV: Understanding and Learning KV Cache Eviction in Full-Duplex Speech Models

Abstract

Full-duplex spoken dialogue models listen and speak at once, and every audio frame of both speakers appends key-value states to the cache, so its memory grows with elapsed session time, even during silence, rather than only with the amount of speech exchanged. We first analyze what this cache holds on two full-duplex models, MiniCPM-o 4.5 and Raon-SpeechChat. The cost of losing a token depends on its content and on whether the token belongs to the user or to the agent, which neither recency nor accumulated attention measures. Eviction under a budget can also leave the model taking the turn while the reply loses its content, a failure that the next-token loss tracks poorly because the loss scores a given reply token by token rather than the reply that the model generates under eviction. We propose DuplexKV, a lightweight -Gate trained jointly with the backbone under a fixed cache budget with an -regularized objective and combined with attention and recency at inference, together with Judged Eviction Tuning (JET), which tunes the gate on an LLM judge's verdicts about the replies the model generates. On the user interruption task of Full-Duplex-Bench, DuplexKV with JET attains the highest judge score among prior eviction policies on both backbones even at small budgets, responding to almost every interruption, and its advantage grows as conversations become longer.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.