SparseSAM: Structured Sparsification of Activations in Segment Anything Models
Abstract
The Segment Anything Model (SAM) achieves strong promptable segmentation, but its ViT-based image encoders dominate inference latency and memory footprint. Existing activation-compression methods, such as token merging, shorten the processed sequence but introduce non-trivial runtime overhead and can suffer severe quality degradation under high compression. Sparse-attention methods typically leave the MLP dense, limiting end-to-end speedup. We propose SparseSAM, a training-free structured-sparsification framework that jointly accelerates attention and MLP layers while preserving token identity. SparseSAM introduces Stripe-Sort Attention, which forms fixed spatial groups using Z-order, ranks those groups with an image-dependent saliency permutation computed once per image, and applies a fixed A-shaped sparse pattern compiled into the attention kernel. SparseSAM further introduces a Residual-Consistency MLP that routes only informative tokens through the MLP while propagating the remaining tokens through the residual pathway. On SAM-L , attention-only SparseSAM reaches latency speedup at 25% density with an absolute mAP decrease of 0.8 mAP points, while joint attention–MLP compression reaches latency speedup with an absolute decrease of 2.7 mAP points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.