Norm or Direction? Decoding Vision Mambas for High-Resolution Vision
Abstract
Vision Mamba models replace quadratic self-attention with linear-complexity selective state space models (SSMs). MambaOut, however, shows that a Gated CNN block can match or exceed VMamba on image classification, questioning whether SSMs are necessary for vision. We ask whether the two models nonetheless encode visual information differently. Using centered kernel alignment (CKA), we find that VMamba's final stage diverges from both MambaOut and its own preceding blocks, and we trace this divergence to two components. The token mixer sets how far information spreads. VMamba's SS2D combines local channels with channels whose memory spans the entire image, some of which remain global at without retraining. The classifier head sets which part of a token the prediction reads. MambaOut pools before normalizing and concentrates class evidence in high-norm foreground tokens, whereas VMamba normalizes before pooling, places high-norm tokens in the background, and preserves class evidence in token directions. Swapping only the head order moves VMamba's high-norm tokens to the foreground. Consistent with direction-based evidence aggregating more stably as tokens grow, VMamba spreads logit support more broadly, degrades less under resolution transfer across three model scales, and outperforms MambaOut on segmentation after full fine-tuning. We suggest token mixing range, token magnitude and direction, and the normalization–pooling order as design axes for high-resolution and dense prediction backbones.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.