DuplexKV: Understanding and Learning KV Cache Eviction in Full-Duplex Speech Models
Abstract
Full-duplex spoken dialogue models listen and speak at once, and every audio frame of both speakers appends key-value states to the cache, so its memory grows with elapsed session time, even during silence, rather than only with the amount of speech exchanged. We first analyze what this cache holds on two full-duplex models, MiniCPM-o 4.5 and Raon-SpeechChat. The cost of losing a token depends on its content and on whether the token belongs to the user or to the agent, which neither recency nor accumulated attention measures. Eviction under a budget can also leave the model taking the turn while the reply loses its content, a failure that the next-token loss tracks poorly because the loss scores a given reply token by token rather than the reply that the model generates under eviction. We propose DuplexKV, a lightweight -Gate trained jointly with the backbone under a fixed cache budget with an -regularized objective and combined with attention and recency at inference, together with Judged Eviction Tuning (JET), which tunes the gate on an LLM judge's verdicts about the replies the model generates. On the user interruption task of Full-Duplex-Bench, DuplexKV with JET attains the highest judge score among prior eviction policies on both backbones even at small budgets, responding to almost every interruption, and its advantage grows as conversations become longer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.