acceptodds
Under review as a conference paper at ICLR 2027

MLP-Bench: Towards End-to-End Formal Research Assistance in Machine Learning Theory

Abstract

Large language models (LLMs) have demonstrated strong mathematical reasoning and autonomous research abilities, yet whether they can assist the rigorous construction and verification of machine learning (ML) theory remains underexplored. Since ML theory rests on precise assumptions and derivations, we ground our study in Lean, which mechanically checks every derivation against its stated assumptions. Mirroring how theory is actually developed and refereed, we frame formal ML theory research as two complementary capabilities: *forward construction*, which assists theory derivation, and *backward verification*, which assists reviewing. We instantiate this framework in the **M**achine Learning **L**ean **P**roof **Bench**mark (**MLP-Bench**), a 360-problem benchmark spanning 16 topics. *Forward construction* is realized by two tasks, formalizing natural-language (NL) theorems into Lean statements and proving given Lean statements. *Backward verification* is realized by a third task, judging whether a proof is correct and, if not, localizing the first erroneous step with Lean-checked evidence. Evaluating 11 models, *forward construction* remains challenging. Auto-formalization reaches at most a **27.5%** bidirectional equivalence (BEq) pass rate against gold Lean statements, with omitted implicit assumptions a major failure mode, and specialized provers solve **0.0%** of theorem-proving problems. In *backward verification*, models localize errors more reliably by constructing Lean-checked refutation than by attempting to prove each step. To our knowledge, MLP-Bench is the first Lean benchmark for end-to-end LLM assistance in ML theory research.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.