RePair: Rubric-Guided Turn-Level Preference Learning for Tool-Using Agents
Abstract
Tool-using agents on long-horizon tasks keep making the same decision errors, and the two standard remedies miss them for different reasons. Supervised fine-tuning imitates expert trajectories, which never visit the states the learner reaches after its own mistakes. Reinforcement learning from a terminal reward does visit those states, but it pays for full rollouts and still cannot say which turn went wrong. We introduce RePair (Rubric-guided Repair as preference Pairs), which turns the learner's recurring errors into a small set of reusable behavioral rubrics. A rubric names an observable error pattern, states the rule that corrects it, and gives a criterion that a repaired turn must meet. Whenever a rubric matches a failed turn, a repair model rewrites the turn as the rule prescribes, and a validator checks the rewrite against the criterion using only what the agent could see at that turn, so no rollout is needed. The rewrite and the learner's original response then form a preference pair on the very same input, and since that input carries no rubric, the trained policy needs none at inference. On BrowseComp-Plus, rubrics induced automatically from the learner's own failures lift a supervised Qwen3-14B search agent from 42.6% to 53.2% with 2,097 pairs. The learner can also write its own repairs: with no teacher supplying the preferred responses, two rounds of self-repair still gain 7.2 points over the same baseline, and the second round adds to the first, so an agent can keep improving from its own failures.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.