RepScan: Representation-Adaptive Scanning for Visual State Space Models
Abstract
High-resolution visual representation learning is hindered by the quadratic complexity of Vision Transformers and the rigid scanning strategies used by existing efficient architectures, which employ uniform token processing or fixed scanning orders that fail to adapt to the highly non-uniform distribution of semantic information in real-world images. To address this, we propose RepScan, a representation-adaptive visual scanning framework that dynamically constructs input-dependent serialization orders. RepScan first introduces Representation Relation Modeling to characterize directed token relations by jointly considering representation affinity and predictive discrepancy. Based on these relations, Adaptive Scan Initialization selects a content-dependent starting token with strong outgoing relational support, while Adaptive Scan Ordering determines the subsequent traversal according to the remaining representation structure. A differentiable permutation relaxation further enables end-to-end learning of the discrete scanning process. RepScan preserves the efficient state-space propagation of the underlying visual SSM while introducing a sparse, input-dependent scan-construction process for adaptive serialization. Extensive experiments across classification, detection, and segmentation demonstrate that RepScan achieves a favorable accuracy–efficiency trade-off across diverse visual recognition tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.