acceptodds
Under review as a conference paper at ICLR 2027

Which Layer Runs the Program? Normalization, Not Depth, Decides Where a Transformer Executes Each Step

Abstract

When a transformer is trained from random initialization to execute a Turing-complete instruction set, it reaches the input–output behavior of a hand-designed reference circuit, while realizing it through a different internal mechanism TwoMechanisms2026, NandaEtAl2023. That comparison fixes which algorithm gradient descent learns; it leaves open a second, equally basic question: where, across the network's depth, is each step of that algorithm computed, and what decides the placement? We study this question on the same substrate as similar recent work, a four-layer transformer trained to execute SUBLEQ Mavaddat1988, using the same hand-built circuit , now as a per-layer answer key: for every datapath signal (operand fetch, dereference, difference, branch, program-counter update), fixes the layer at which it could be computed, and we require every probe to recover that layout before trusting it. Read this way, the trained network reproduces 's pipeline in the same dependency order. Where that pipeline runs, however, is decided by an unexpected component: normalization. In a paired, multi-seed comparison whose only difference is a pre-normalization layer, the dereference and the subtraction it feeds are deferred to the final layer without normalization and, with it, staged early at the layer uses, under both LayerNorm and RMSNorm. The shift holds on a target no affine function of the operands can supply, so it is not an artifact of when the operands become readable. We give an idealized-model account of the effect and turn it into a control. A short analysis shows a block's gradient scales with the norm of its input, which normalization holds to a fixed value (up to a learned gain) at every depth. Acting on this, we place normalization at a chosen layer to relocate the computation to that depth, show the contrast persists across depth, the normalized share decreasing toward , and find it dissolves under a norm-free replacement. As by-products we localize that stage to attention rather than the feed-forward block and show the effect recurs across one-instruction-set variants and on modular addition.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.