acceptodds
Under review as a conference paper at ICLR 2027

Pushing the Math Reasoning Frontier by Making Beyond-Capability Problems Learnable

Abstract

Large language models (LLMs) have made substantial progress in mathematical reasoning, which is largely driven by reinforcement learning (RL). However, further pushing the mathematical reasoning frontier of LLMs through RL remains challenging, as models rarely generate successful rollouts on problems beyond their current capability frontier, resulting in a scarcity of effective learning signals. To address this challenge, we propose **LeanTree**, which decomposes beyond-capability problems into learnable subproblems. The major challenges in developing LeanTree are twofold. (1) LLM-based decompositions rely on stronger models and still introduce logical hallucinations. (2) The resulting subproblems lack a difficulty hierarchy, leaving the training order unclear. LeanTree leverages the intermediate reasoning structure of existing formal proofs (machine-verifiable representations of mathematical reasoning) to generate *reliable* and *hierarchical* subproblems. By training on these hierarchical subproblems, the model progressively acquires the reasoning skills needed to solve the corresponding beyond-capability problems. Extensive experiments on mathematical reasoning benchmarks demonstrate that LeanTree provides more effective learning signals and improves overall reasoning accuracy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.