acceptodds
Under review as a conference paper at ICLR 2027

When Every Rollout Fails: Splicing Learner-Conditioned Options into Long-Horizon Agent RL

Abstract

Group-based reinforcement learning with verifiable rewards often yields all-fail groups on long-horizon agent tasks: every sampled trajectory receives zero reward, leaving group-relative methods without a reward-based policy gradient. A common remedy supplies complete solutions from a stronger model, but full trajectories are expensive and expose the learner to long off-policy histories. We introduce SPLICE, which constructs local training targets from failed current-policy trajectories online. For each all-fail group, a proposal model selects a splice point in a current-policy trajectory and generates several bounded options from that point. We estimate each option's value using repeated current-policy continuations, convert positive estimates into a return-weighted local target, and fit the current policy on the options. We evaluate SPLICE across two policy families on two long-horizon agent tasks, Software Engineering with SWE-bench Verified and Deep Research with BrowseComp-Plus. SPLICE outperforms GRPO and full-trajectory guidance baselines in all four settings, improving over GRPO by 2.2–10.8 percentage points. On initially failed SWE tasks, it raises GRPO's graduation rate by 10.7 percentage points. Further experiments show that our method improves performance while reducing costly proposal-model computation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.