MathIdeaBench: A Benchmark For Mathematical Research Idea Generation
Abstract
As large language models (LLMs) continue to advance in mathematical reasoning and progress from problem solving to research discovery, rigorously evaluating their capacity for research ideation becomes vital. However, measuring the quality of mathematical research proposals remains an open challenge due to its technical and subjective nature. In this work, we introduce MathIdeaBench, the first benchmark designed to measure the quality of LLM-generated mathematical research proposals against a human baseline. MathIdeaBench consists of (1) an automated data generation pipeline which produces a refreshable evaluation dataset, and (2) an LLM-based evaluator, MathIdeaJudge, which scores proposals using detailed evaluation rubrics designed by experts. Furthermore, we propose a tree-based diversity metric for proposals based on mathematical subject classification. Our experiments show that compared to ideas from human-authored publications, LLMs tend to generate proposals which are less sound, less feasible and less diverse but conceptually deeper and more original.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.