acceptodds
Under review as a conference paper at ICLR 2027

ResceneKV: Reconstruction-Error Selective KV Cache Eviction across Nested Scales for Visual Autoregressive Modeling

Abstract

Visual autoregressive (VAR) modeling generates images by progressively predicting token maps from coarse to fine, through a next-scale paradigm that enables efficient sampling and strong zero-shot generalization. However, predicting each scale from all preceding scales requires retaining the corresponding keys and values in the KV cache. As generation proceeds to finer scales, the KV cache therefore grows continuously and becomes a major memory bottleneck for high-resolution image synthesis. Existing KV cache eviction methods typically rely on attention statistics to guide eviction decisions, although eviction candidates with similar attention mass can exhibit substantially different reconstruction errors. To address this mismatch, we propose ResceneKV, which directly measures the output reconstruction error induced by KV cache eviction and uses it for joint KV cache budget allocation. We further exploit the redundancy between the conditional and unconditional branches of classifier-free guidance by reusing the conditional KV cache for the unconditional branch at later scales. Extensive experiments across image generation benchmarks demonstrate that ResceneKV consistently improves pixel-level fidelity over existing KV cache eviction methods while maintaining generation quality across different memory budgets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.