Entropy as a Lens: Revealing Structural Importance in Masked Self-Supervised Vision Transformers
Abstract
Masked self-supervised Vision Transformers (ViT) such as MAE and VideoMAE have demonstrated strong transferability but often exhibit substantial architectural redundancy. In this work, we ask a simple yet underexplored question: Are all Transformer blocks equally important for representation learning? We present a data-free analysis framework that estimates block importance using the entropy of pretrained weight-number, without relying on downstream data or task-specific sensitivity measurements. Through mathematical derivation and empirical analysis, we show that block-wise weight entropy strongly correlates with performance degradation when blocks are removed, revealing a clear hierarchy of importance across Transformer depth. Leveraging this observation, we demonstrate that a large portion of blocks in masked self-supervised ViT can be removed with minimal impact on downstream performance, suggesting significant structural redundancy in pretrained models. Beyond compression, our results provide new insights into how representation capacity is distributed across Transformer blocks, and position entropy as a lightweight diagnostic tool for understanding and analyzing large pretrained vision models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.