acceptodds
Under review as a conference paper at ICLR 2027

Twilight: Thinking with Generated Images via Lightweight KV Cache

Abstract

Thinking with Generated Images (TwGI) has emerged as a promising paradigm that enables vision-language models to actively synthesize novel visual content as intermediate reasoning steps, transcending the limitations of static input images. However, due to the inclusion of numerous image tokens and multiple classifier-free guidance (CFG) samples, TWGI's reasoning process incurs significant memory pressure, limiting scalability to long-sequence generation. By analyzing attention patterns, we reveal that TwGI's KV cache contains significant redundancy, as some heads maintain low attention scores on previous tokens. Therefore, we propose Twilight, the first training-free KV cache eviction framework tailored for TwGI. Specifically, we introduce a novel normalized standard deviation metric to identify modality-specific retrieval heads that attend globally to historical tokens. Leveraging the distinctive attention patterns of retrieval heads and streaming heads across text and image generation stages, we design targeted KV cache eviction strategies and sparse attention mechanisms. Our approach selectively retains critical tokens while discarding redundant ones, significantly reducing memory consumption without compromising generation quality. Extensive experiments validate the effectiveness of our approach. Twilight maintains or even improves generation performance through a training-free approach, while reducing KV-cache memory by over 60% on average.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.