Beyond Episodes: Learning Behaviors That No Rollout Demonstrated
Abstract
An LLM agent’s rollouts can contain complementary fragments that together solve a task, even when no individual rollout succeeds. Single-episode supervision leaves these combinations undemonstrated. We investigate whether agents can learn these behaviors by composing their own fragmented interaction experience. We propose GraphStitch, which constructs goal-conditioned training trajectories by joining fragments from different rollouts at shared states. By pooling both successful and failed attempts, it recovers solutions absent from every original rollout. Specifically, GraphStitch organizes observed transitions into a shared graph, extracts composed paths, and assigns goals using environment-specific achievement rules. State identity or environment replay checks executability before ordinary supervised fine-tuning. This extends supervision beyond episode boundaries while reusing recorded action-labeled edges as its building blocks. We evaluate compositional goal reaching and generalization to held-out tasks. On held-out Blocksworld, GraphStitch improves greedy success over hindsight relabeling by 28.2 percentage points. Using one-quarter as many collected rollouts as Hindsight, GraphStitch also achieves higher mean held-out greedy success on Blocksworld (68.0% versus 46.8%), with additional replay during data construction. Gains also appear on Countdown, and controlled comparisons support the contribution of cross-episode composition. Composing fragmented experience thus enables LLM agents to learn beyond behaviors demonstrated in individual rollouts.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.