acceptodds
Under review as a conference paper at ICLR 2027

Reinforced Generative Recursive Reasoning for Native Adaptive Computation Depth

Abstract

Recursive Reasoning Models scale computation by iteratively refining a latent state with shared transition functions, yet the amount of recursion is usually fixed externally or predicted by an auxiliary head. We show how a generative recursive prior turns this limitation into a learnable decision: its Gaussian transition is a normalized, reparameterizable policy with an exact log-density. Variational pretraining therefore provides an on-manifold starting distribution, while outcome reinforcement learning can update both latent transitions and a residual-conditioned halting action. A potential-based value signal supplies credit across the loop and a KL anchor keeps exploration near the pretrained manifold. On structured reasoning benchmarks, ReGRAM improves ARC-AGI-1/2 accuracy by 4.8/5.7 percentage points over GRAM, allocates more depth to harder inputs, and transfers from a training depth cap of 16 to inference budgets of 1024. The approach motivates a simple principle: adaptive depth is most natural when the computation itself is represented as a stochastic policy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.