acceptodds
Under review as a conference paper at ICLR 2027

Some Tokens Think, Others Remember: Separating Compute and Capacity in LLMs

Abstract

Looped transformers apply a shared block multiple times and have emerged as a parameter-efficient route to scaling compute in language models. However, at fixed FLOPs a looped model has strictly less capacity than a baseline transformer. We argue that compute, the number of sequential operations applied to a hidden state, and capacity, the parameters available at a single step, are distinct scaling axes that standard transformer layers conflate. We propose a dual-path block that exposes both axes as parallel pathways within a single layer: a deep sublayer re-applied 𝐾 times with shared parameters, and a wide sublayer with an enlarged feed-forward network applied once, combined by independent per-token gates. At two FLOP budgets, the best dual-path configuration surpasses both iso-FLOP single-axis controls on aggregate language modelling, commonsense, and math metrics, while using 17–33% fewer parameters than the width-scaled baseline. The learned gates are directly interpretable and show systematic per-token allocation: function words and lexical content trend wide, while punctuation, symbols, and arithmetic tokens trend deep. Our results suggest that compute and capacity can be scaled separately, with the model balancing the two token by token.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.