acceptodds
Under review as a conference paper at ICLR 2027

Code2Math: Can Your Code Agent Evolve Math Problems Through Exploration?

Abstract

While large language models have achieved remarkable progress in solving mathematical problems, reaching IMO-level performance and tackling research-level tasks, the ability to formulate new and more challenging mathematical problems remains an important yet underexplored capability. This capability is particularly relevant to recursive self-improvement, where models are expected to autonomously expand the scope and difficulty of the tasks they can engage with. Meanwhile, recent code agents provide executable environments for symbolic computation, systematic search, and empirical verification, making them a natural substrate for autonomous mathematical exploration. To this end, we introduce Code2Math, a multi-agent system that leverages code agents to evolve mathematical problems through executable test-time exploration. Our experiments demonstrate that, given sufficient test-time exploration, code agents can synthesize new, solvable problems that are structurally distinct from and more challenging than the originals. Moreover, augmenting the original training data with the corresponding evolved versions of these problems further improves model performance compared with training on the original data alone, demonstrating the practical value of problem evolution for model training. Together, these results demonstrate the potential of code-driven problem evolution as a scalable mechanism for autonomous mathematical exploration and a building block toward self-evolving reasoning systems.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.