xRRMs: Explorative Recurrent Reasoning Models
Abstract
Recurrent Reasoning Models (RRMs), including Hierarchical Reasoning Models (HRM) and Tiny Recursive Models (TRM), achieve strong performance on combinatorial reasoning tasks such as Sudoku while remaining parameter efficient. HRM and TRM are in principle looped transformers in which a shared module is applied repeatedly to refine a neural state. Conceptually, this computation is grounded in a fixed-point equation, using deep supervision and stop-gradients between iterations. Recently, stochastic variants of RRMs, such as Equilibrium Reasoners (EqR) and Probabilistic Tiny Recursive Models (PTRM), have been proposed. Beyond enabling generative modeling, stochasticity provides a natural mechanism for improving exploration of the solution space. In particular, these variants induce stochasticity by randomizing the initial latent states and by injecting noise during recurrent iterations. However, their gains rely on exploring multiple trajectories at inference, which raises the question of how to encourage exploration more directly. We address this with xRRM, a Recurrent Reasoning Model trained with strategies from Explorative Modeling: multiple sampled neural states are refined in parallel. At each deep supervision step, the loss is evaluated for all sampled trajectories, but gradients are computed only from the trajectory whose output is closest to the correct solution. This effectively implements a winner-takes-all objective and restricts backpropagation to the best-matching trajectory. On Sudoku-Extreme, xRRM achieves an average fully solved rate of more than 90% with a single evaluation candidate, improving over EqR by more than 20 percentage points. Test-time scaling with 32 evaluation candidates increases this value to more than 99%. We extend our approach to generative problem settings and evaluate it on the Game-of-24 benchmark, where multiple valid solutions per problem exist, investigating whether effective exploration extends beyond deterministic problem solving.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.