acceptodds
Under review as a conference paper at ICLR 2027

Training Recursive Language Models for Hierarchical Generalization

Abstract

Language models may excel at applying familiar reasoning patterns yet struggle when those patterns must be combined in a novel way. Recursive inference offers a natural way to tackle novel compositions by decomposing them into potentially familiar subproblems and composing their results. Prior work has largely studied recursion for long-context processing and efficient inference, leaving its effect on compositional generalization unclear. We introduce a controlled Countdown benchmark for hierarchical compositional generalization and compare recursive and linearized execution on the same reasoning trajectories with matched training and inference token budgets. Recursive inference improves accuracy on held-out compositions up to while generating about 40% fewer tokens. Separated subproblem contexts also permit training on reasoning trajectories whose total length greatly exceeds one context window, and extending search may further help by exploring alternative ways to combine these familiar constituents. Training on wider recursive trajectories and allowing wider search at inference increases compositional accuracy by , whereas doubling the generation limit yields only a 10% relative increase. Allocating search budgets across recursion levels improves the accuracy–compute trade-off, while mixing search budgets during training achieves similar quality with roughly half as many generated tokens. Finally, we build an adaptive curriculum by utilizing the compositional relationships between the tasks, doubling the held-out compositional accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.