Towards Lean-grounded Mathematical Reasoning
Abstract
Formal verification and informal mathematical reasoning in LLMs have largely evolved independently. Informal outcome-level supervision provides limited guidance on how rigorous proofs should be structured, while invoking proof assistants such as Lean at inference time incurs recurring formalization costs. To address this, we propose Lean-grounded Learning, which uses Lean as privileged training-time supervision to improve informal reasoning without requiring Lean at inference. A teacher first generates an informal proof, then revises it using a verified Lean proof. A student is subsequently trained on these Lean-grounded revisions via SFT. On Qwen3-30B-A3B, Lean grounding improves AIME24 accuracy from 45.0 to 55.0 and ProofNet from 72.1 to 80.9 over the baseline with natural-language privileged information. With only 30K Lean-grounded examples, Qwen2.5-Math-7B-Base surpasses Qwen2.5-Math-7B-Instruct on MATH, AMC, and Olympiad Math, despite the latter being trained on roughly 2.5M examples followed by RL. Experiments show that Lean-grounded models more often derive claims from established facts, whereas Qwen3-30B-A3B more often relies on candidate testing and extrapolation from concrete cases.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.