Mamba-B3: Block-Structured, Bounded, and Bijective Frequency SSM for Vision
Abstract
Vision state-space models (SSMs) typically use real-valued diagonal transitions, which efficiently propagate information but cannot represent rotational dynamics that naturally arise in periodic textures, oriented structures, and motion. We introduce Mamba-B3, a vision SSM that combines three complementary modifications to the state transition dynamics. B1 replaces independent real-valued channels with block-structured decay–rotation dynamics and introduces low-rank cross-frequency coupling within each head, allowing frequency components to interact during state evolution. B2 controls the resulting non-normal dynamics through a spectral budget and a block-wise reachability regularizer, providing a sufficient condition for stable state propagation while improving input conditioning. B3 uses bilinear discretization, which, within the one-parameter family of methods, is uniquely able to preserve imaginary-axis modes on the unit circle while maintaining a one-to-one mapping between continuous rotation rates and discrete angles. We evaluate Mamba-B3 on ImageNet-1K, ADE20K, COCO, Something-Something V2, and Breakfast. Across image classification, dense prediction, and video recognition, Mamba-B3 consistently improves over the corresponding vision-Mamba baselines, while systematic component and interaction ablations support the roles of coupling, stability conditioning, and discretization. These results suggest that explicitly modeling and controlling frequency dynamics is a promising direction for more expressive visual SSMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.