acceptodds
Under review as a conference paper at ICLR 2027

ReMath: Benchmarking Theorem Retrieval in Research-Level Mathematics

Abstract

Mathematical proofs often use known theorems. We introduce ReMath, a bench- mark for retrieving theorems needed in unfinished research proofs. Given the global setup and a local proof prefix, a large language model (LLM) uses external search to retrieve a theorem statement with its source. The source lets researchers check the statement and credit prior work. The retrieved theorem must justify the intended next step using the available facts. The task allows suitable alternative theorems and sources. We build a dataset from published papers across five mathe- matical domains. Across the tested models, most outputs that pass the grounding check do not pass the reference-based sufficiency check. Case studies show that a theorem may assume a fact the proof still needs to establish, while a conclusion weaker than the reference may suffice for the next step. Evaluation must therefore check how the theorem applies to the intended step, including whether the theorem concerns the objects in the proof. We release the dataset and code for construction and evaluation

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.