CandiKV: Preserving Candidate KV State Across Selection and Continuation in LLM Workflows
Abstract
Branching LLM workflows score several candidate histories before continuing from a selected subset. Scoring has already materialized the KV state of every candidate, yet request-scoped serving separates that state from the continuation that may consume it. When selection resolves, the chosen histories return as new requests, and the engine must recover work it has just performed before their first continuation tokens. We introduce CandiKV, which treats selection as a handoff of candidate state and keeps the candidate group live until the choice is known. While selection is pending, incoming scores guide computation and KV residency toward likely continuations. After the choice, selected histories enter generation with their state and progress intact. On six reasoning and retrieval workloads, CandiKV reduces continuation TTFT by up to 90.7%. It also cuts the GPU KV byte-seconds wasted on unselected candidates by 23.7–30.5% compared with keeping every candidate resident, while regular-request TTFT changes by at most 0.33%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.