acceptodds
Under review as a conference paper at ICLR 2027

Ramanujan Challenge Results: AI-Assisted Evaluation Pipeline for Multi-Format Proofs

Abstract

The Ramanujan Challenge put AI-assisted mathematics proofs to the test, inviting the community to prove previously unpublished formulae involving mathematical constants using AI. The number, mathematical level, and diversity of submissions presented a substantial challenge: assessing the correctness of complex AI-assisted proofs at scale. Treating the Challenge as a research-level math benchmark, we introduce an AI-assisted pipeline to measure submission correctness, solution similarity, and the success of different AI models, Computer Algebra Systems (CAS), and formal tools in mathematical proofs. We introduce an AI-assisted evaluation pipeline for diverse mathematical evidence, including natural language proofs, Computer Algebra code, and Lean formalizations. Beyond simple execution or compilation, it evaluates semantic and logical correctness by verifying assumptions, bounds, and completeness of statements for Computer Algebra and Lean code. Comparing to a one shot prompting judgment setup, we show a reduction in false positive verdicts (31.25% to 3.75%) for natural language proofs, and an improvement in judgment accuracy (37.8% to 56.8%) for supporting code. Using this pipeline, we find that participants - some without research-level math experience - were able to generate correct submissions for every problem, including solutions conceptually distinct from the official proofs, and solve one of the published open conjectures. Together, the benchmark and evaluation framework provide a systematic methodology for assessing AI-assisted research mathematics at scale.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.