Janus: Breaking Layer-Bound KV-Cache Layouts to Accelerate Local Agent Inference
Abstract
As local agent conversations grow longer, the key-value (KV)-cache can exceed commodity GPU memory, making KV-cache offloading necessary but introducing substantial host-to-device (H2D) transfer latency. Sparse self-speculative decoding (SSP) is a promising way to reduce transfer latency because the sparse KV-cache can be small enough to fit on GPU during drafting, while the full KV-cache is fetched only to verify multiple candidate tokens together. However, the straightforward combination of SSP and offloading (naive SSP) incurs redundant H2D transfers and layout rematerialization. This is because existing layer-bound layouts cannot accommodate the different KV-cache requirements of draft and verification, preventing reuse of overlapping KV-pages. We propose Janus, a local agent inference system whose layer-decoupled KV-page abstraction manages physical KV-pages in a unified GPU pool to enable cross-phase reuse. To keep large-scale metadata lookups and updates efficient, a coalesced metadata manager processes them through batched and fused primitives. To use the limited GPU page budget efficiently, a hierarchical GPU KV-page allocator first selects the verification staging depth and then calibrates the sparse draft budget offline. We implement Janus on vLLM V1 and evaluate three LLMs on two commodity GPUs across three agent workloads. Compared with baselines, Janus reduces time per output token (TPOT) by 26.2%-78.0% and end-to-end latency by 18.8%-61.8%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.