Beyond Linear Scaling: Recurrent Calibration for Cross-Resolution Transfer in Pretrained Visual Mamba Models
Abstract
Mamba has emerged as a leading paradigm for long-sequence modeling and has been widely extended to vision. Despite its linear complexity and appeal for high-resolution representations, pretrained visual Mamba models often degrade severely at unseen resolutions. In this paper, we reveal that resolution scaling in visual Mamba is not merely a one-dimensional sequence extension, but entails a fundamental change in the two-dimensional spatial rate of pretrained selective dynamics. The resulting structural mismatch propagates through state-dependent residual updates and hierarchical transitions. To quantify this resolution-induced trajectory drift, we formulate Cross-Resolution Trajectory Discrepancy (CRTD). Pre-calibration analysis reveals a strong association between raw trajectory discrepancy and cross-resolution accuracy degradation. Based on these insights, we propose Architecture-Native Uniform Residual/Update Calibration (Uniform-RU), a lightweight and label-free framework that calibrates recurrent dynamics and residual trajectories using only a small fraction of the training images while keeping the pretrained backbone frozen. Across diverse visual Mamba architectures and resolutions, Uniform-RU substantially restores cross-resolution accuracy on image classification and generalizes well to detection and segmentation tasks. These results demonstrate that recurrent calibration can effectively extend pretrained visual Mamba models, offering a dynamics-centric perspective for resolution transfer beyond sequence scaling. Code is available at https://anonymous.4open.science/r/Uniform-RU/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.