-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Labeled QA Data
Abstract
Deep search agents remain difficult to train at scale due to limited labeled data and weak credit assignment in long-horizon search. Self-play offers a scalable way to reduce data dependence, but conventional approaches optimize models only with sparse outcome rewards, limiting learning efficiency. We find that the task-generation process in self-play can itself serve as a scalable source of privileged supervision for self-distillation. Specifically, self-play naturally produces a Question Construction Path (QCP), which captures the reverse solution process and provides rich intermediate signals for self-distillation without curated privileged information. Leveraging this insight, we propose Privileged Information Self-Play (-Play), a self-evolution framework that couples self-play with self-distillation by directly exploiting privileged information generated during task construction. In -Play, self-play generates training tasks together with QCPs for self-distillation in a low-cost and scalable manner, while self-distillation converts QCPs into token-level supervision to accelerate self-play evolution. This closed-loop interaction enables more efficient model improvement without relying on labeled QA data or curated privileged information. Across multiple benchmarks and model scales, data-free -Play surpasses fully supervised search agents and improves evolutionary efficiency by 1.5–3.1 over conventional self-play.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.