ClueFlow: Learning Long-Horizon Search through Attributed Research-State Transitions
Abstract
Long-horizon search requires agents to acquire relevant evidence and integrate it into an evolving research state. Existing methods maintain research states to guide exploration or assign turn-level progress signals, yet grounding these signals in explicit, attributable state changes remains underexplored. We introduce ClueFlow, which learns long-horizon search from attributed research-state transitions. ClueFlow constructs Source–State–Final attribution chains that connect evidence-acquiring browser actions, Evidence and Candidate updates, and the final answer supported by those updates. During search, ClueFlow records typed state changes and potentially supporting browser actions, then applies gold-conditioned semantic validation and source tracing after execution. We use this attribution to guide both supervised fine-tuning and reinforcement learning. During SFT, attribution selects which decisions receive supervision with their original interaction contexts. During Turn-PPO, the same attribution augments outcome-based turn advantages with credit assigned along the Source–State–Final chains. Transition-level validation also recovers useful supervision from gold-consistent state changes in trajectories with incorrect final answers. With Qwen3.5-35B-A3B, ClueFlow achieves 59.1% accuracy on BrowseComp and 72.0% on BrowseComp-Plus. In controlled BrowseComp-Plus experiments with Qwen3-30B-A3B, attribution-guided SFT matches all-turn exact-context SFT using only about 30% of the training tokens, while attribution-guided PPO outperforms outcome-only PPO by 4.1 percentage points. These results demonstrate that state-transition attribution enables efficient supervision and effective credit assignment for long-horizon search agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.