acceptodds
Under review as a conference paper at ICLR 2027

Do Hierarchical Vision Backbones Need a Fourth Stage?

Abstract

Hierarchical vision backbones devote substantial capacity to their final stage. We investigate how much of this capacity is needed and how learning without it changes the preceding hierarchy. We train Swin and ConvNeXt models from Nano to Base with and without the fourth-stage blocks, always keeping the final downsampling transition. Small and Base models trained without these blocks use about 29% fewer parameters and 25% less peak inference memory at batch size one, and lose only 0.13–0.53 ImageNet-1K Top-1 points. Reallocating this capacity upstream can improve accuracy at comparable parameter counts, at the cost of more FLOPs, whereas the narrowest configurations still favor four stages. Under a common normalized readout, class information becomes accessible earlier within the third stage in all ten model pairs, and truncated models generally depend more on the late spatial exchanges of this stage. On Swin-T-C64, a terminal token-mixing module with 8.6% of the original stage's parameters recovers about half of the truncation loss and outperforms a parameter-matched channel-only control. COCO box and mask AP stay within one point of the full models in the evaluated configurations, and ImageNet-C gaps are modest at Small and Base scale but larger in some smaller configurations. These findings support choosing terminal capacity jointly with upstream allocation and the evaluation objective.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.