acceptodds
Under review as a conference paper at ICLR 2027

Endogenous Signals Fail to Allocate: Per-Token Compute Value Is Real, Transient, and Not Learnable

Abstract

Token-level adaptive compute usually keys off endogenous signals (attention entropy, router scores, confidence) assumed to track per-token difficulty. On arithmetic with graded difficulty we show this assumption fails three ways: an entropy-driven gate collapses to a constant, a tag-initialized gate inverts, and an unsupervised learned head is vacuous. Difficulty-supervised targets under a budget constraint do produce ordered gates, but a pre-registered shuffled-targets control shows the ordering carries no accuracy benefit at 11M (at most 0.3 accuracy points), and the 1.4B replication differs from permuted targets by 0.0008 held-out loss. Measuring the field directly, the per-token gradient of loss with respect to each token’s compute gate, explains why: the field has a life cycle (real structure mid-training, consumed by training, flat at convergence), and a causal test shows its structure is not trainable: chasing it online loses to random allocation of the same budget at every scale tested. The surviving recipe needs no allocator: token-level FFN dropout, where each token receives full FFN compute only inside randomly rotating windows. At 1.4B web-text fine-tuning it trails dense by 7.79% held-out loss (3 seeds) at 50% billed FFN training FLOPs, inside the pre-registered 2–10% band; a gather-based realization converts half the billed cost into a measured 1.35× training throughput. Failures on code and template-dominated math are reported.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.