acceptodds
Under review as a conference paper at ICLR 2027

ViS: Vision State-Space Token Mixer

Abstract

State-space models (SSMs) bring linear-complexity sequence modeling to vision, but only after the image is flattened into a 1D sequence under a handcrafted scan order, a causal bias that multi-directional scanning mitigates rather than removes. In contrast, we view SSMs from the perspective of token mixing: each token can aggregate information from other tokens by iteratively updating its own hidden state, as in an SSM recurrence. This perspective suggests that what must be defined is not a single global scan order, but a principled order in which each token aggregates its context. We introduce the Vision State-Space Token Mixer (ViS), which redefines the scan axis as spatial distance: every token's state absorbs information from concentric rings around it, yielding a scan-order-free token mixer with a global receptive field. Despite the apparent difficulty of parallelizing such position-wise recurrences, we derive an exact closed-form formulation as structured masked attention, enabling efficient implementation on modern attention-friendly hardware. Experiments across image classification, segmentation, object detection, and generation demonstrate ViS as a viable general-purpose vision backbone; ViS achieves up to higher throughput while maintaining comparable or better performance, and exhibits the strongest zero-shot resolution transfer among vision backbones through its coordinate-computed spatial prior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.