Variable-Width Transformers
Abstract
Scaling model size, specifically depth and width, has driven significant progress in transformer-based language models. However, most architectures maintain a constant width across layers, allocating a fixed parameter and computation budget despite their potentially different computational roles. We empirically investigate nonuniform capacity allocation by proposing a ×-shaped > <former architecture. It maintains wider early and late layers while narrowing the middle layers, using a parameter-free residual resizing mechanism. Across decoder-only language models ranging from 200M to 2B parameters (dense) and 3B parameters (MoE), our > <former consistently outperforms parameter-matched uniform baselines on language modeling loss. By reducing the average layer width, this architecture also requires fewer overall FLOPs (22% reduction under fitted loss-matched scaling curves) and smaller KV cache memory and I/O cost (15% reduction). Consequently, dense > <formers reach the uniform baseline’s final smoothed training loss with 15.6–23.1% fewer profiled GPU-hours, and generation throughput improves by 8.8% in one profiled setting. In analysis, we find qualitatively different residual-stream representations. Overall, nonuniform width allocation can result in more resource-optimal scaling of language models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.