Simple approach to token pruning for vision state-space models
Abstract
Vision Mamba models offer linear sequence complexity but remain costly at token counts of high resolution images. Token pruning can reduce this cost, but methods developed for Vision Transformers do not directly account for the recurrent, order-dependent computation of state-space models. Removing tokens changes subsequent state updates and can disrupt information propagation along the sequence. We introduce DeltA, a parameter-free token-pruning importance metric designed for vision state-space models. DeltA scores each token using the interaction between its input-dependent timescale parameter and the magnitude of its post-block representation. The method uses signals already available in the backbone and requires no learned pruning module. Experiments on image classification, semantic segmentation, and facial-attribute recognition using ViM and PlainMamba architectures show a favorable accuracy–computation trade-off and further indicate that the same criterion generalizes across different vision Mamba architectures and tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.