acceptodds
Under review as a conference paper at ICLR 2027

SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

Abstract

Looped Transformers iterate a shared block of layers to gain effective depth without adding parameters, and prior work reports that this extra depth improves language modeling, multi-hop reasoning, and math. However, most of these evaluations compare models of the same size, where the looped model also spends more FLOPs per token, so its gains mix architectural advantage with extra compute. Meanwhile, work that compares Looped Transformers with unlooped ones at the same FLOPs often finds that the looped models perform worse, but there the unlooped models store more parameters. Therefore, to investigate whether looping itself has an advantage, we study looping on Mixture-of-Experts Transformers and compare each looped model with an unlooped Baseline that closely matches it in per-token FLOPs, total non-embedding parameters, and KV cache. Through three ablations, we arrive at an effective recipe we call SMELT (Sparse MoE Transformer, middle layers Loop Twice), which loops the middle half of layers twice and pays for the extra layer executions by narrowing the model. We scale SMELT and the Baseline up to 54B total non-embedding parameters and fit a separate Chinchilla-style scaling law to each architecture. SMELT's loss drops faster with compute, saving 14.7–18.0% of training compute on the compute-optimal frontier at FLOPs. Beyond validation loss, SMELT's advantage exceeds what its loss gap predicts on downstream benchmarks and grows with sample length and the number of in-context examples. Mechanistic analysis shows that, on their second pass, the looped layers put less attention on the attention sink and more on content-relevant tokens, an inductive bias that may underlie the observed performance gains. These results show that looping can improve Transformers even under matched budgets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.