AREX: Verification-Guided Recursive Refinement for Deep Research
Abstract
Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This asymmetry suggests that agents can use the verification results of a provisional answer to direct subsequent research toward the constraints that remain unresolved. We introduce AREX, a verification-guided deep research agent that recursively refines its research state and answer. Its inner research loop gathers evidence, maintains structured records containing source URLs and supporting summaries, and returns a provisional answer, an evidence list, and an answer-level confidence score. The outer refinement loop uses this confidence to accept the result or, when confidence is insufficient, assesses the current trajectory and either launches a targeted refinement round or restarts the investigation. To support long-horizon research, AREX learns an autonomous context-update tool that consolidates its interaction history without relying on an external model. We train dense 4B and 122B-A10B MoE variants on verified synthetic tasks and high-quality trajectories, emphasizing decision-critical steps during both mid-training and reinforcement learning. Under a unified maximum inference budget, AREX achieves strong performance across six research and reasoning benchmarks, while additional recursive rounds yield consistent test-time scaling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.