State-Space Models as Vision Encoders for VLMs: A Study across Training Settings and Interfaces
Abstract
Vision-Language Models (VLMs) typically leverage a vision encoder whose image features are mapped into a Large Language Model through a connector. While transformer-based vision encoders remain the standard choice, we ask whether state-space models (SSMs) can provide a viable alternative, given their competitive performance on vision tasks. To answer this question, we evaluate VMamba across four pre-trained setups (i.e., contrastive, classification, detection adaptation, and segmentation adaptation) and find that, under a shared downstream VLM recipe, the resulting models are generally competitive across VQA and referring-expression benchmarks, with the most consistent gains on grounding tasks. Moreover, we investigate the vision-language interface and find that (i) adding a normalization layer before the connector can substantially improve the VLM performance, and (ii) increasing the depth of an MLP connector produces inconsistent gains, while representation probes show that different connector depths alter the balance of local and global information in the visual tokens. Together, these results position VMamba as a competitive vision encoder alternative for VLMs, and highlight the importance of controlling the connector interface when comparing frozen vision encoders.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.