acceptodds
Under review as a conference paper at ICLR 2027

Pretrained Mamba-Based Vision Backbones Can Exhibit Collapsed Timescale Banks: Checkpoint-Aware State Expansion with SpanInit

Abstract

Expanding the recurrent state of a pretrained Mamba requires initializing the decay timescale of each added state dimension. Additional-scan, a parameter-efficient state-expansion method, copies a pretrained entry of the state matrix , so each added state begins at a decay scale already present in that channel's timescale bank. We show that the effect of this initialization varies with the timescale geometry inherited from pretraining. Because within-channel timescale ratios depend only on , the bank's relative span can be read directly from a checkpoint without data or a forward pass. Across released Mamba-based vision checkpoints, ImageNet-1k-only models with multiple state dimensions per channel exhibit near-collapsed timescale banks, whereas models exposed to ImageNet-21k have substantially wider spans. Motivated by this diagnosis, we propose SpanInit, which fills the largest uncovered gaps within the scan-relevant range in log-timescale space. SpanInit changes only the initial values and uses a single calibration forward pass to set absolute timescales, adding no trainable parameters or inference cost. On VTAB-1k, SpanInit improves Additional-scan on the ImageNet-21k-pretrained MambaVision-B/L checkpoint but degrades it on the ImageNet-1k-only checkpoint of the same architecture. These contrasting outcomes indicate that state-expansion initialization should account for the pretrained checkpoint's timescale geometry rather than follow a single checkpoint-agnostic rule.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.