SPD: Skeleton-Preserving Distillation from Quadratic to Linear-Time Vision Models
Abstract
The accuracy of Vision Transformers relies on large-scale pretraining with a quadratic-cost attention mixer. Subquadratic vision backbones avoid this cost at inference but are usually pretrained from scratch. We propose Skeleton-Preserving Distillation (SPD), which converts a pretrained ViT into a linear-time model by reusing its pretrained weights: the student inherits every component of the teacher except the quadratic mixer. We apply SPD to MambaFormer, our ViT-based backbone with a bidirectional Mamba2 mixer in place of attention, distilled from ViT teachers whose classifier path it can inherit. Using only 10% of ImageNet-1K, MambaFormer-L recovers 98% of its ViT-L teacher's accuracy at 18.6× less compute than re-training the teacher, and 99.7% with full data. With a CLIP ViT-B/16 teacher, MambaFormer-B reaches 84.25 top-1 vs. 84.30 for ViT-Linearizer at 5.5× less compute, and reaches 82.26 from only 10% of the data at 20.6× less. Trained from scratch on labels with 3.0× more compute (teacher pretraining excluded), the same model reaches 67.75. SPD also applies to semantic segmentation, where the student stays within 0.5 mIoU of the teacher on ADE20K and Cityscapes (single-scale evaluation), beats a no-teacher baseline by 2.1–2.4 mIoU, and needs 1.4× fewer inference FLOPs on Cityscapes. With fused scan kernels, MambaFormer-B matches ViT-B's throughput at 224² (batch 32) and is 1.55× faster at 1200² on an H100.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.