ProSearch: Teaching LLMs When (Not) to Search over Streaming Inputs via Reinforcement Learning
Abstract
We introduce ProSearch, a reinforcement-learning framework for training proactive search agents for streaming inputs. Given a continuous input stream, the agent must decide whether to act, what external knowledge to retrieve, and how to respond, under streaming latency constraints. Existing systems either retrieve on every input or delegate the trigger to an external classifier, leaving the “whether to act” decision outside the learned policy. ProSearch instead internalizes it, training a single policy end-to-end with GRPO so that the trigger, query, and grounded answer are emitted by the same model and optimized against the same reward. The central challenge is reward sparsity: actionable segments are rare and correct behavior requires a three-step pipeline (trigger, retrieve, respond), so an outcome-only reward yields degenerate gradients. We address this with a structured reward decomposition whose three terms each score a different sub-event of the rollout, densifying the learning signal for GRPO relative to an outcome-only reward and yielding stable training. Across three proactive tasks spanning claim verification and knowledge-seeking turn detection, ProSearch consistently beats reactive and externally-orchestrated baselines, with up to a 2 gain in end-to-end task success.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.