SeamKV: Efficient Construction of Full-Prompt KV Factors under Chunked Prefill
Abstract
Long-context LLM inference faces growing KV-cache memory and long-prompt prefill. KV-cache compression reduces resident state, while chunked prefill processes prompts incrementally; they compose poorly when the compact representation itself must be built from the full prompt. In xKV-style cross-layer low-rank sharing, each factor depends on all prompt tokens but is produced by a local layer group. Under chunk-major prefill, factor completion is tied to the final chunk, causing dense K/V from multiple groups to overlap near prompt end. We present SeamKV, which instead sends all prompt chunks through one factor-producing group before advancing deeper. It accumulates factor statistics during forward execution, completes remaining prompt-global products over bounded source blocks, then replaces dense K/V with the finished factor after the group has produced all required rows and its last prefill use has passed, preserving the same full-prompt factorization target. Relative to a matched chunk-major control, SeamKV reduces per-GPU peak memory by 9.5–20.8% across 32K–128K, while cache-ready time is lower at every tested length; at 128K, the peak number of simultaneously live dense groups falls from 8 to 1. Controlled experiments isolate earlier producer completion as the source of the memory reduction; in the matched causal audit, factor-finalization CUDA time differs by at most 0.33%. The schedules compute the same full-prompt operator in exact arithmetic, and under BF16 their paired RULER and LongBench results remain closely matched. The effect transfers to a per-layer full-prompt key-factor geometry, reducing peak memory by 13.6–22.6%; all 120 paired examples produce identical token sequences under SeamKV and the matched control. Prompt-global factorization therefore need not imply prompt-end construction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.