acceptodds
Under review as a conference paper at ICLR 2027

Train on the States You Deploy: Closed-Loop Hard Routing

Abstract

A 31M learned sparse-attention model with native perplexity reaches perplexity when deployed with the hard routes it was trained to produce. Soft-only training omits the hard deployment-state distribution; because deployed routes are generated layer by layer, that distribution is closed-loop. We introduce *Closed-Loop Hard Routing* (CLHR), which keeps a differentiable soft path for router learning while training the backbone on the hard trajectory induced by its own decisions. In an exact reproduction of a public 31M-parameter protocol, CLHR reduces closed-loop deployment excess from to nats across three seeds, changing the hard/native perplexity ratio from to . With the configuration frozen, the reduction transfers to 300M models on WikiText-103 () and FineWeb-Edu (); on LAMBADA, standard hardening reduces accuracy from to , whereas CLHR changes it from to . A matched causal ladder shows that hard exposure, open-loop replay, and shuffled routing do not suffice. Across the tested settings, successful soft-to-hard deployment requires task supervision on useful, self-induced hard states.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.