acceptodds
Under review as a conference paper at ICLR 2027

Mathematica Incognita: Characterizing Open-Ended Mathematical Exploration in Language Models

Abstract

Mathematics has long depended on problem-solving as a guidepost for structuring and measuring progress—a methodology that is strongly reflected in the development of mathematical AIs. However, open-ended exploration, as well as the related activities of theory building and asking new research questions, remain under-studied capabilities of language models. In order to evaluate open-ended exploration of language models, we develop a suite of problem-solving tasks within equational theories—sets of equational axioms that define an algebraic structure. This area of math is particularly useful for model evaluation due to the ease of instantiating diverse, self-contained theories that are rarely described in literature, as well as the possibility of generating problems within these theories which remain difficult for frontier reasoning models without external tool use. Our experimental tasks evaluate two desiderata of successful explorations: future utility and idea diversity. As a proxy for future utility, we measure the improvement in problem-solving performance when models are provided with an exploration document, and observe notable but uneven benefits from exploration across models and theories. To measure idea diversity, we implement a limited automated reasoning query interface for theory explorations that enables unambiguous comparison of research ideas across model runs, and observe a common core and long tail of ideas across models. As language models gain a greater role in solving research-grade math problems, studying their behavior during open-ended exploration will become increasingly important for comprehensive understanding of their research capabilities.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.