acceptodds
Under review as a conference paper at ICLR 2027

Search-E1: Self-Distillation Drives Self-Evolution in Search-Augmented Reasoning

Abstract

Outcome based reinforcement learning provides a simple and effective recipe for training search augmented reasoning agents. Recent work seeks further improvements through stronger external teachers, auxiliary process supervision, specialized reward shaping, and increasingly elaborate training pipelines. This raises a natural question: do search agents really need this additional machinery to continue improving? We present Search-E1, a simple self evolution framework that alternates vanilla GRPO with on policy self distillation (OPSD). After each GRPO round, Search-E1 mines sibling trajectories generated by the policy itself and exposes the search skeleton of an efficient successful trajectory, that is, the sequence of queries it issued, as privileged context. A token level forward KL objective then aligns the policy under its standard inference context with its own distribution under this privileged context. In this way, the policy converts contrasts among its own rollouts into dense token level supervision without an external teacher, auxiliary reward model, or additional annotation. On seven single hop and multi hop QA benchmarks, Search-E1 reaches an average Exact Match of with Qwen2.5-3B and with Qwen2.5-7B, achieving the best average performance among all compared methods at both scales. These results show that alternating outcome based exploration with self distillation provides a simple and effective path to iterative improvement in search augmented reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.