Inlay: Accelerating Position-Independent KV Caching with Sparse Repairs and Shared Decode Kernels
Abstract
Position-independent KV caching (PIC) reduces repeated computation by reusing cached text chunks across different contexts. However, this reuse does not automatically yield compact KV storage or efficient attention execution during decoding. Request-specific repairs can lead to duplicated KV states, while request-oriented attention repeatedly loads the same shared data. We present Inlay, a serving runtime that combines compact KV storage with shared attention execution for PIC decoding. Inlay separates immutable shared Base KV, token-granular request- and layer-specific Repair KV, and private append-only Decode KV without changing the underlying PIC method’s repair policy. Its attention kernels reuse loaded Base KV across compatible queries, apply request-specific repair masks and position adjustments, and merge partial attention outputs. A cost-aware CPU planner selects between shared execution and private KV packing for each chunk under a memory budget. Valid plans, metadata, and packed prompt KV are reused across decode steps to reduce preparation overhead. On NVIDIA A100 attention microbenchmarks using Qwen3-8B and Llama-3.1-8B configurations, Inlay achieves median kernel speedups of 1.24× over FlashInfer and 1.32× over PAT across the tested workloads and attention geometries. In single-layer integrations with CacheBlend, EPIC, and QCFuse, Inlay achieves median kernel speedups of 1.43–1.98× and median resident KV memory reductions of 50.1–63.6% relative to their request-private FlashInfer baselines, while preserving the same logical inputs and PIC repair results. Resident KV measurements exclude additional packed execution copies.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.