acceptodds
Under review as a conference paper at ICLR 2027

RAHL: Recursive Self-Improvement through Repairability-Aware Hindsight Learning

Abstract

Language agents can fail their assigned tasks while achieving other goals along the way. Hindsight relabeling can recover these successes, but changing the goal alone may leave recorded plans inconsistent with the new task. Goals achieved by the same trajectory can differ in how much plan revision they require. We introduce Repairability-Aware Hindsight Learning (RAHL), which selects verified goals according to the cost of repairing their associated plans. The acting policy revises plans in execution order, conditioning each edit on the previously repaired history. A policy-generated mask restricts edits to spans classified as goal-specific intentions, while recorded actions and observations remain unchanged. Among candidates that satisfy the editing constraints, RAHL favors those requiring fewer token changes. Each replay-verified demonstration is paired with new rollouts from the same initial state for training alongside original tasks. The updated policy then acts, proposes goals, and repairs demonstrations in the next round, forming a recursive actor–relabeler loop. Experiments show higher GRPO task success on ALFWorld, WebShop, and ScienceWorld at matched training rounds, with gains depending on the optimizer. On identical WebShop histories, the trained relabeler yields more replay-verified demonstrations, while demonstration yield per model call remains similar.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.