acceptodds
Under review as a conference paper at ICLR 2027

Mixture-of-Powers: Learning to Fast-Forward Looped Transformers

Abstract

A conventional looped transformer advances computation by repeatedly applying the same learned update. Accelerating that process requires a shortcut whose output remains usable by subsequent updates. Mixture-of-Powers (MoP) trains separate experts for strides of two, four, eight, and sixteen base updates. An externally supplied schedule combines experts and unit steps through a shared hidden state. Training supervises intermediate boards and execution after a jump, while keeping the base model, encoders, and readout fixed. On Sudoku with a prescribed sequence of forced placements, one stride-16 call reaches the requested board on 98.24% of 512 test windows, compared with 99.22% for sixteen base calls. It executes 32 rather than 64 transformer blocks and reduces single-example end-to-end time from 54.49 to 27.25 ms. Across 25 call sequences excluded from training, mean endpoint exactness is 98.34%. Longer paths expose a limit: four stride-16 calls reach 97.85% endpoint exactness, but only 81.05% match every scheduled intermediate board. These single-seed results show an inference benefit for a prescribed recurrent computation, with additional parameter and training costs and imperfect intermediate-state fidelity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.