acceptodds
Under review as a conference paper at ICLR 2027

Tiny Recursive Models Underuse their Depth

Abstract

Tiny recursive models (TRMs) solve hard reasoning benchmarks by repeatedly applying a small weight-tied network to a pair of carried states, and improvements to this family have so far come from iterating more: extra recursion cycles, extra supervision steps, more effective depth. However, a failure case emerges: the recursion reaches a fixed point quickly. Successive iterations change the carried states less and less, so much of the depth budget is spent after the recursion has effectively reached its fixed point, and the added depth is ineffective. We show that a larger residual stream alleviates this failure case. Adopting Virtual Width Networks (VWN), we widen the state the recursion carries several-fold while leaving backbone compute unchanged, so each iteration has room to write something the previous one did not and the fixed point becomes harder to reach. Widening also makes the depth TRMs already run useful: held-out sweeps show the wide model's margin over the narrow one growing with the number of supervision steps allowed at evaluation when many candidate answers are scored, while the narrow model's candidate pool stops improving early. We validate on ARC-AGI-1, ARC-AGI-2, and Sudoku-Extreme with 7M-parameter models under matched configurations. The resulting VWN+TRM improves ARC-AGI-1 pass@2 by at least points over both TRM baselines we trained (the upstream and the depth-matched configuration), improves ARC-AGI-2 pass@2 by points over the matched baseline ( points, or relative, over the published TRM result), and adds points of exact accuracy on Sudoku-Extreme.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.