acceptodds
Under review as a conference paper at ICLR 2027

LURE: Pursuit-Evasion Self-Play for Zero-Data Multi-Step Reasoning

Abstract

Reinforcement learning with verifiable rewards (RLVR) has substantially improved large language model reasoning, but its scalability remains tied to externally supplied tasks whose difficulty must keep pace with an improving policy. Zero-data self-play addresses this bottleneck by generating its own curriculum, yet existing approaches either estimate task difficulty from model-derived proxies or rely on executable task specifications, while multi-step solvers remain largely trained from sparse outcome supervision. We introduce LURE, a verifier-grounded pursuit-evasion framework for zero-data self-play, in which a task-generating Challenger and a planner-executor Solver co-evolve through a shared environment verifier. This formulation turns the verifier into a common learning interface for both curriculum construction and policy improvement. Specifically, LURE uses verifier-confirmed Solver success to position generated tasks near a capture frontier, and converts stepwise verifier progress into capture-anchored dense process credit for the Solver. Round-anchored regularization further stabilizes the two adaptive policies during co-evolution. Across three heterogeneous reasoning environments and three backbone families, LURE outperforms advanced baselines under both unified and specialist settings, and across out-of-domain benchmarks it delivers the strongest near-domain transfer and the best overall average.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.