acceptodds
Under review as a conference paper at ICLR 2027

From Plausible Repairs to Policy Capability

Abstract

A language model can propose a plausible correction for a failed attempt without knowing whether it will help the agent finish. Even when the intervention works, deciding what the policy should learn from it is a separate problem. We introduce RePair, a research harness that connects failure analysis to policy learning. RePair restores the decision state where a proposed change applies and compares the baseline and modified continuations under matched conditions, with the same frozen policy completing both branches. Only interventions that repeatedly improve task completion become positive training evidence. Using the same verified evidence and action coverage, we then compare action-only supervision with a dual-view alternative that also teaches the subgoal, expected progress, and recovery conditions behind the action. Both train separate descendants of the same parent policy, while deployment retains the original action-only interface. Finally, we remove external failure memory and intervention to test whether any improvement remains in the policy itself. On ALFWorld with Qwen2.5-3B-Instruct, we study the effects of verification and training representation together with the cost of finding useful corrections. RePair makes the path from plausible correction to policy capability experimentally testable by separating what is worth learning from how that evidence is retained in the policy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.