acceptodds
Under review as a conference paper at ICLR 2027

Re:Zero: Learning from Failed Trajectories through Verified Behavioral Recomposition

Abstract

Failure-only rollout groups represent a supervision dead end for group-based agentic RL: they offer no successful demonstrations, and under binary rewards their group-relative advantages vanish. Yet these failed trajectories often harbor complementary, useful sub-behaviors. We introduce Re:Zero, a framework that converts all-fail groups into verifiable supervision without additional policy decoding. Given an all-fail group, Re:Zero replays and simplifies failed trajectories, then forms bounded prefix–fragment recompositions by splicing a source prefix with a contiguous donor fragment across different positions. A completion-aware search evaluates candidates within a fixed interaction budget, retaining solely executions that reproduce native environment success (or satisfy a strict score threshold). The policy then learns continuation actions grounded in the replayed observation history, masking the shared prefix to prevent credit confusion. Experiments across ALFWorld, WebShop, ScienceWorld, and Visual Sokoban with models from 1.7B to 27B demonstrate that Re:Zero consistently improves base group-RL algorithms at zero extra policy decoding cost during training, preserving the base deployment pipeline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.