acceptodds
Under review as a conference paper at ICLR 2027

ResearchMath-14K: Scaling Research-Level Mathematics Via Agents

Abstract

The frontier of mathematics is defined by problems whose solutions are not yet known. However, whether language models can meaningfully engage with such problems without human intervention remains unclear. A major obstacle is the lack of large-scale research-level math datasets. To this end, we introduce ResearchMath-14k, a set of problems curated from academic sources via a multi-agent pipeline. ResearchMath-14k spans 11 mathematical domains and ranks above existing math datasets on knowledge, novelty, and procedural difficulty. To our knowledge, it is the largest research-level mathematical problem set available for training. We additionally generate K teacher trajectories and use prompting and behavioral filtering. Notably, however, generating correct trajectories is nontrivial at this level, and two LLM judges label only and of sampled ResearchMath training trajectories as correct. Nevertheless, across three model families, full-parameter training on ResearchMath improves performance on graduate- and research-level mathematics benchmarks by points over the starting checkpoints. In comparison, training on existing datasets such as DASD and Nemotron-SFT-Math-v4 changes performance by and points, respectively. Notably, mixing DASD with ResearchMath yields higher scores than token-matched DASD alone on benchmarks covering olympiad short-form (), graduate- and research-level short-form (), graduate- and research-level symbolic (), and proof evaluation (). Further analysis suggests that research-level mathematical content and greater reasoning diversity may help explain these gains. These broad gains indicate that ResearchMath provides complementary supervision to contemporary datasets. We make publicly available for future works on research-level mathematical reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.