acceptodds
Under review as a conference paper at ICLR 2027

Tapered Language Models

Abstract

Modern language models, including transformer, recurrent, and memory-based variants, share a common chassis: a stack of identical layers in which parameters are allocated uniformly across depth. This default is inherited from the original transformer and has remained largely unchanged, yet a growing body of evidence suggests that layers contribute non-uniformly to the final output. We ask whether parameter capacity should reflect this asymmetry. In a controlled experiment at a fixed parameter budget, allocating more MLP capacity to earlier layers and less to later layers improves perplexity over a uniform-width baseline, while the reverse allocation, the direction taken by prior layer-wise scaling schemes, hurts. Building on this result, we propose *Tapered Language Models* (TLMs), in which MLP width decreases monotonically across depth under a fixed total budget. We focus on MLPs because they dominate parameter count across modern LM families and expose width as a single, clean axis of variation. Using a smooth cosine schedule selected once on a M Transformer and transferred unchanged, tapering improves average downstream accuracy and perplexity over matched uniform baselines in all six comparisons across three architectures (Transformer, Gated Attention, and Titans) at M and B parameters, with identical parameter count and FLOPs. These results suggest depth-aware capacity allocation as a simple, architecture-agnostic guideline for language model design.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.