Pushing the Math Reasoning Frontier by Making Beyond-Capability Problems Learnable
Abstract
Large language models (LLMs) have made substantial progress in mathematical reasoning, which is largely driven by reinforcement learning (RL). However, further pushing the mathematical reasoning frontier of LLMs through RL remains challenging, as models rarely generate successful rollouts on problems beyond their current capability frontier, resulting in a scarcity of effective learning signals. To address this challenge, we propose **LeanTree**, which decomposes beyond-capability problems into learnable subproblems. The major challenges in developing LeanTree are twofold. (1) LLM-based decompositions rely on stronger models and still introduce logical hallucinations. (2) The resulting subproblems lack a difficulty hierarchy, leaving the training order unclear. LeanTree leverages the intermediate reasoning structure of existing formal proofs (machine-verifiable representations of mathematical reasoning) to generate *reliable* and *hierarchical* subproblems. By training on these hierarchical subproblems, the model progressively acquires the reasoning skills needed to solve the corresponding beyond-capability problems. Extensive experiments on mathematical reasoning benchmarks demonstrate that LeanTree provides more effective learning signals and improves overall reasoning accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.