Is Depth All You Need? Rethinking ViT Capacity via Structural Re-parameterization
Abstract
Vision Transformers (ViTs) are computationally expensive largely due to their deeply stacked layer designs, yet prevailing efficiency research has concentrated on algorithmic-level strategies such as token pruning and attention mechanism acceleration. This raises a fundamental but underexplored architectural question: is it possible to decrease ViT depth without sacrificing model capacity? We address this question by introducing a branch-based structural re-parameterization framework that augments transformer blocks with parallel branches during training, thereby enhancing optimization flexibility and model capacity. After training, these branches are systematically consolidated into a single linear pathway through exact re-parameterization applied at the inputs of nonlinear activations, ensuring strict functional equivalence between the expanded and collapsed models. Experiments across multiple ViT backbones demonstrate that this approach achieves a 2–4 reduction in network depth while maintaining competitive classification performance on ImageNet-1K. Beyond depth reduction, we further demonstrate that re-parameterization functions as a general capacity enhancer, boosting the accuracy of iso-depth ViTs and improving vision encoders within vision-language models (VLMs). These findings indicate that heavily stacked transformer architectures may be redundant for error-tolerant and resource-constrained deployment scenarios, opening promising avenues for the design of efficient vision transformers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.