acceptodds
Under review as a conference paper at ICLR 2027

VTDS: A Vision Semantic Time-Delay State Space Mixer for Efficient Long-Range Visual Modeling

Abstract

Visual state space models (SSMs) provide an efficient alternative to self-attention for long-range visual representation learning. However, existing approaches mainly improve state transition dynamics or spatial scan paths, while the content-dependent semantic distance between a token and its relevant context remains implicit. To address this limitation, we propose VTDS, a Visual Semantic Time-Delay State Space Architecture, which explicitly incorporates token-wise semantic delay modeling into visual state propagation. VTDS learns a differentiable dependency distribution over a bounded historical window to retrieve semantically relevant context, and adaptively integrates the retrieved information with the recurrent state through a bounded fuzzy regulator. A loop-free diagonal recurrence further enables parallel state propagation. For a fixed lag window, VTDS scales linearly with sequence length without constructing a global token-wise attention matrix. Experiments on ImageNet-1K classification, COCO object detection and instance segmentation, and ADE20K semantic segmentation demonstrate competitive accuracy-efficiency trade-offs across diverse visual recognition tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.