Knowing When to Stop: IGP-Bench and the Multi-Agent System GALOIS for the Inverse Galois Problem
Abstract
Frontier agents can now sustain hours of autonomous work. Greater capability may bring greater progress, but it may also carry an agent farther along a wrong route. Existing long-horizon mathematics benchmarks are hard to build, saturate quickly, or grade partial progress only with a language model, which makes it difficult to reliably measure how much effort an agent spent on each route it tried. We introduce IGP-Bench, a benchmark of long-horizon algebraic geometry problems in inverse Galois theory. Its first release has 200 problems in three tiers: 39 with published explicit realizations, 55 whose realizations are known to exist but have not been made explicit, and 106 that remain open. On every problem, a successful route consists of objects that are hard to find but easy to verify, and GAP and Magma check each of them deterministically. Agents record their trajectory as an artifact graph, from which we score verified progress, format compliance, tool use, and knowing-when-to-stop metrics: a useful-resource fraction and a stopping-efficiency score comparing how far an agent goes down incorrect routes with how far it goes down correct ones. We test eight frontier models over 80 runs on the five Mathieu-group tasks, each run with a 10-hour wall-clock limit and 500,000 output tokens. Together these tasks offer 32,979 candidate routes, of which only 13 lead to a realization. Surprisingly, we find that Claude Opus 5 has the highest mean progress (76.2), followed by GPT-6 Sol (62.1) and GPT-6 Astra (59.4), and that the gap between them is not one of construction ability but of identifying the right route and of where the budget goes. These gaps motivate GALOIS, an open-source multi-agent system with a lead agent, a strategy reviewer, and specialist sub-agents for its 36 written skills, in which a dedicated route rejection agent screens candidate routes for known obstructions in the background, so wrong routes are ruled out quickly. In reference runs for IGP-Bench, GALOIS with a Claude Opus 5.5 lead produces Magma-verified polynomials for all five Mathieu tasks in 10 to 126 minutes, including and , for which no single agent produced a polynomial. Tasks, verifiers, evaluation code, and GALOIS will be publicly released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.