VTDS: A Vision Semantic Time-Delay State Space Mixer for Efficient Long-Range Visual Modeling
Abstract
Visual state space models (SSMs) provide an efficient alternative to self-attention for long-range visual representation learning. However, existing approaches mainly improve state transition dynamics or spatial scan paths, while the content-dependent semantic distance between a token and its relevant context remains implicit. To address this limitation, we propose VTDS, a Visual Semantic Time-Delay State Space Architecture, which explicitly incorporates token-wise semantic delay modeling into visual state propagation. VTDS learns a differentiable dependency distribution over a bounded historical window to retrieve semantically relevant context, and adaptively integrates the retrieved information with the recurrent state through a bounded fuzzy regulator. A loop-free diagonal recurrence further enables parallel state propagation. For a fixed lag window, VTDS scales linearly with sequence length without constructing a global token-wise attention matrix. Experiments on ImageNet-1K classification, COCO object detection and instance segmentation, and ADE20K semantic segmentation demonstrate competitive accuracy-efficiency trade-offs across diverse visual recognition tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.