acceptodds
Under review as a conference paper at ICLR 2027

How Long Does a Model Need to See? When Vision Becomes State

Abstract

Diffusion vision-language models (DVLMs) keep the image in the computation throughout denoising, tying seeing to answer formation by design. We ask a simple question: how long does a model actually need to see? By terminating visual computation at different stages, we find that seeing and thinking have different lifetimes: visually grounded answers can continue to form after the image has left the computation. Visual influence changes substantially over denoising and can re-emerge after becoming weak, while visual attention largely misses this transition. Meanwhile, visual information becomes increasingly reflected in the evolving generation state, whose structure reveals whether vision will matter again later. This signal is surprisingly simple to read out: a 305-parameter predictor can forecast future visual influence, determine when visual computation can end, and transfer across DVLMs. It preserves approximately full task performance while substantially reducing decoding attention and visual KV usage. Together, these results reveal a simple visual lifecycle in diffusion multimodal generation: seeing can end once vision becomes state.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.