Same Conversation, Different Tokens: Characterizing and Mitigating Token Drift for Efficient Multi-Turn LLM Serving
Abstract
Multi-turn LLM serving reuses cached KV states only when a new prompt extends the token sequence the engine has already processed. Applications, however, carry conversation history as text: rendering it through a chat template and encoding the result again can return the same conversation as a different token sequence, silently defeating prefix reuse. We call this divergence token drift and characterize it across six open models, three conversation datasets, and four languages. Re-encoding alone breaks the prefix in 4.6%–90.8% of requests at temperature 1.0, at a roughly constant rate per generated token, while template history rewriting breaks it on every request after the first in three of the six model configurations. Since reuse ends at the first mismatch, rebuilding history costs 4×–23× the prefill the new content alone would need, and the KV branches stranded by misses add 10%–122% to the live context's footprint while remaining indistinguishable from it under recency-based eviction. We design one technique per layer of the serving pipeline and integrate all three into vLLM: canonical generation at the sampler with CG-Recheck, a checker we design that constrains sampling to token sequences that survive re-encoding; token-level history (TLH) at the request interface, which carries earlier turns forward as the token IDs the engine produced; and DriftEvict, a KV-cache eviction policy that identifies a dead branch when a turn completes and reclaims it ahead of live context. The two preventive techniques raise cache hit rates where their layer can prevent the miss, and TLH improves time to first token (TTFT) the most; but both change what the model samples or sees and may lower output quality. DriftEvict leaves outputs unchanged and improves throughput and TTFT under high load on every evaluated model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.