Learning Recursive Reasoners without Solutions
Abstract
Recursive reasoning models obtain effective depth by repeatedly applying small networks, but are typically trained on ground-truth solutions. We show that such models can instead be learned without solution supervision, using only a black-box verifier or energy function that scores candidate solutions. Our solution-free variant combines latent deterministic recurrence with stochastic refinement of explicit candidate assignments and is trained from terminal verifier feedback alone. On constraint satisfaction and combinatorial optimization problems, such as Graph Coloring, Maximum Independent Set and Sudoku, the resulting models are effective solvers, matching the state of the art. To our knowledge, we provide the first solution-free model to exceed 90% on Sudoku-Extreme. We further find that latent recurrence is a test-time scaling axis complementary to explicit refinement and parallel search, and that the compute-optimal allocation across these axes is problem-dependent. Finally, our results show that more training-time recurrence does not necessarily improve depth extrapolation, whereas truncating credit assignment to the final recursive cycle yields operators that keep improving well beyond their training depth.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.