Width Expansion via Parallel Block Interaction for Vision Model Training
Abstract
Increasing the width of a vision backbone improves its representation capacity. However, directly widening increases the cost of dense channel processing. This paper argues that block-level lateral connectivity, a dimension largely abandoned since the advent of deep learning, deserves to be reopened. We propose Parallel Block Interaction (PBI), a general block-level width expansion framework that distributes the feature channels across multiple parallel blocks while enabling lightweight information exchange among them. Each parallel block retains the original backbone-specific block design while processing a subset of the channels. The shared public channels provide an internal information path across parallel blocks without changing the feature width passed to downstream modules. Public-to-public, private-to-public, and private-to-private channel interactions enable sufficient and controllable information exchange across blocks. On ConvNeXt, ViT, Swin Transformer, and VMamba backbones, PBI improves image classification and dense prediction performance under comparable parameter and computation budgets. Comparisons with direct channel expansion and Mixture-of-Experts, and full ablation studies show favorable accuracy–efficiency trade-offs and clear contribution of each component. Class-level and spatial-region analysis further validates that PBI strengthens the representation diversity of the parallel blocks. PBI reveals the advantage of width expansion via block-level intra-layer connectivity, a potential model scaling dimension orthogonal to “going deep”.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.