FastPIC: Fast Position Independent Caching for Multi-Tenant LLM Serving
Abstract
As LLM context windows continue to expand, prefill overhead has become major bottlenecks in long-context serving. Existing context-caching systems rely on prefix matching and cannot effectively reuse repeated content that appears at different positions across multi-tenant requests. Position-Independent Caching (PIC) supports non-prefix reuse, but it remains limited by boundary shifts caused by fixed chunking and attention deviations caused by cross-context KV reuse. This paper presents FastPIC, a position independent KV-cache reuse mechanism for multi-tenant LLM serving. FastPIC uses Token-Defined Chunking to derive boundaries and cache identities from token content, reducing the impact of prefix variations on matching identical token sequences. It further uses Dual-Path Token Selection to recompute attention sinks and high KV-deviation tokens under a limited budget, improving both cache reuse and generation quality. Experiments across multiple models and datasets show that FastPIC reduces TTFT and improves serving throughput while preserving generation quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.