Context-Orchestrated Research Math Agents
Abstract
Long-horizon mathematical reasoning fails less often because a model cannot produce a valid next step than because an agent fails to maintain and expose the right semantic state across many iterations. Left unmanaged, this produces research-level proofs that are locally convincing yet globally incomplete: a key lemma unproved, an assumption unchecked, a citation unsupported, or a computational claim unverified. We present **Research Math Agents (RMA)**, an agentic framework for long-horizon proof development built around a persistent, typed research store and an orchestrator that compiles operation-specific context from that store. The **Research Context Orchestrator** is the central state-management layer between the persistent research store and each locally scoped proof operation: it retrieves task-relevant artifacts, compiles them into a bounded context, invokes the appropriate operation, and writes the resulting proof edits, issue updates, literature notes, plans, or evaluations back to the store. This process is designed to keep proof revisions, unresolved issues, prior attempts, literature, and evaluations available across rounds while exposing only task-relevant state to each local operation. We evaluate RMA across complementary research-level settings using independent expert evaluation, blind mathematician review, LLM-based benchmark evaluation, and Lean 4 kernel verification. RMA achieves a **42.5% solve rate** on the independently evaluated SOOHAK Challenge Hard set, obtains **8 of 10 correct solutions** on First Proof B1 and **8 of 10 passing solutions** on B2 under human-expert evaluation, and verifies **213 of 300** sampled Research Solved targets in Formal Conjectures with the Lean 4 kernel. These results support structured state and context orchestration as useful components for long-horizon proof development, while comparisons against independently developed systems remain subject to differences in inference budgets, backbones, and evaluation protocols. Project page: https://anonymous1776research.github.io/RMA-ICLR2027/
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.