Humanize Olympiad Agents: A Multi-Session Workflow for Olympiad Problem Solving with Solve, Review, Wiki, Verify
Abstract
Large language model (LLM) agents are rapidly improving on long-horizon scientific and olympiad tasks. Recent systems have produced a Lean-formalized resolution of the Navier–Stokes problem and autonomous solutions to four open questions in the Erdos Conjectures database. Yet neither system releases both its underlying model and complete inference stack. More broadly, a single agent session remains brittle on difficult problems: an early error can propagate, missing domain knowledge can stall progress, and the agent can appear to succeed after misreading or weakening the task. We introduce HOA (Humanize Olympiads Agent), an open, flow-level framework that coordinates multiple agent sessions through three complementary feedback channels. First, an independent reviewer checks each candidate and returns actionable feedback. Second, a structured wiki supplies task-specific knowledge that may be missing or unreliable in model pretraining. Third, a verifier checks generated Lean proofs against trusted challenge statements. With GPT-5.6 Sol, HOA reaches the perfect score on IMO, IPhO, and IOI 2026, and near-full score on IChO 2026 problem set and solves all 672 problems (100%) on PutnamBench. HOA solves 216 of 302 problems (71.5%) on Lean-Eval. Furthremore, we extend HOA's results with the open models – With Kimi-K3, HOA also achieves full score on IMO, IPhO,on IOI, and gold-level 97.1% on IChO, which closely matches the closed-source GPT-5.6 Sol + /Goal baseline on IMO and IPhO and exceeds it on IChO and IOI, showing that systematic feedback can compensate for substantial differences in base-model capability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.