VISTA: Learning Wide Search Agents with Verifiable Intermediate States and Turn-Level Credit Assignment
Abstract
Wide search requires agents to explore many target items, track which ones they have covered, and combine their findings. Yet existing systems largely rely on designed multi-agent systems. Training capable open-source wide-search agents is challenging because reinforcement learning with verifiable rewards (RLVR) typically assigns credit to an entire trajectory, making it difficult to identify the few decisions that matter most in long interactions. We propose VISTA (Verifiable Intermediate States for Turn-level Attribution), a reinforcement learning framework that transforms sparse outcome supervision into fine-grained turn-level credit. Our key insight is to make otherwise latent search progress explicit and verifiable through intermediate table states that track the agent's accumulated discoveries. VISTA uses changes across consecutive states to attribute credit to the responsible turns, while preserving the terminal outcome as the dominant learning signal. Across six benchmarks and seven models, ranging from 4B to 27B, VISTA consistently outperforms outcome-only GRPO and step-wise RL baselines. On WideSearch, VISTA improves Item F1 by 13.84% relative over outcome-only GRPO on Qwen3-4B, with consistent gains across other model families and scales, and further generalizes to standard multi-hop QA benchmarks. Extensive ablations show that the gains come not merely from adding process supervision, but from making intermediate progress verifiable and assigning credit to the responsible turns. These results highlight that where and how credit is assigned is critical for training long-horizon search agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.