acceptodds
Under review as a conference paper at ICLR 2027

HorizonKV: Three-State KV Cache Compression under a Shrinking Generation Horizon for Visual Autoregressive Models

Abstract

Visual autoregressive (VAR) KV caches consume massive GPU memory—over 80 GB for Infinity-8B at 1024×1024 resolution. Existing compression relies on binary pruning: each KV state is either kept at full precision or discarded entirely. Our analysis shows that this forced choice is unnecessary: preserving all tokens at lower precision often proves more effective than discarding a fraction of them, demonstrating that fidelity reduction is not information deletion. This insight points to a different path: rather than deciding only whether to keep a state, we should also decide how faithfully to keep it as its future role evolves. To this end, we propose HorizonKV, which replaces binary decisions with a three-state lifecycle Native (full precision), K8V8 (8-bit quantized), and Drop (released) governed by the monotonically shrinking remaining generation horizon. Specifically, a small offline calibration first decouples future demand from compression sensitivity for each state. Afterwards, we allocate a global physical-byte budget via Lagrangian relaxation, enforcing irreversible transitions as future responsibility decreases. Extensive experiments across Infinity-2B and Infinity-8B show that HorizonKV consistently improves the memory–quality frontier: at only 1% KV storage, it approaches full-cache visual quality and achieves roughly 4× additional compression over the state-of-the-art pruning baseline HeatKV at comparable fidelity, while preserving semantic generation quality on GenEval and DPG-Bench.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.