Learning When and What to Teach: Self-Evolving Privileged Search Agents
Abstract
Long-horizon search agents are difficult to optimize because rewards arrive only after many interdependent actions. On hard tasks, all sampled trajectories may fail, eliminating group-relative learning signals and leaving useful recovery actions outside the policy’s support. We observe that search answers are often entities or relations that can serve as training-time anchors for reasoning backward toward better queries. Direct answer conditioning, however, risks privilege leakage rather than transferable search behavior. We introduce MIST-OPD, a single-backbone self-evolution framework requiring no external teacher, language-model judge, or reward model. An answer-blind Student and a privileged Guide are initialized as two trainable policies from the same base model; their asymmetry is information, not capacity. At each public state, the Guide transforms the answer into backward hints at six ordered granularities, ranging from general search strategy to an explicit multi-step route that still withholds the answer. Repeated common-seed rollouts compare no intervention, each candidate granularity, and a source-matched placebo at the same granularity. These matched counterfactuals determine whether guidance is causally useful and select the minimum sufficient intervention for each state. Training alternates updates to both roles. The Guide learns when to abstain, how much information to reveal, and what to teach. The Student learns when to request help and which next action to take through causal token-level OPD, private-token projection, and behavior replay. A lagged Guide stabilizes distillation, and the updated pair generates the next interaction round. At deployment, privileged information and the Guide are removed, leaving only the evolved answer-blind Student. Results show that MIST-OPD achieves state-of-the-art performance, demonstrating its effectiveness as a self-evolving framework for converting privileged answers into transferable long-horizon search behavior without external model supervision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.