acceptodds
Under review as a conference paper at ICLR 2027

Agentic-DPO: Rollout-Free Agent Training from the Student's Own Mistakes

Abstract

Expert trajectories are the cheapest supervision for LLM agents, yet supervised fine-tuning (SFT) never confronts the agent with the mistakes it would itself make at the expert's states. Agentic-DPO mines these mistakes without executing any action during training: at each expert state it samples a few one-step actions from the current student, keeps the most probable one whose decision differs from the expert's, and trains the student to prefer the expert action with a length-scaled DPO loss, stabilized by an SFT anchor and by Policy-Preserving Augmentation (PPA), which re-renders trajectories under different action schemas. We evaluate with ten seeds (five at 32B), validation-selected checkpoints, and twelve tuning configurations for every method. On five closed-loop tasks (τ-bench retail and airline with the standard GPT-4o user, ALFWorld, WebShop, and end-to-end StableToolBench), Agentic-DPO improves over SFT with the same augmentation by 5.9–15.2 points at 9B, matches or exceeds the step-level methods IPR and PivotRL, and recovers 93% (95% CI 77–111) of online GRPO's τ-bench gain with 6.7× fewer GPU-hours; GRPO started from Agentic-DPO has the highest mean on all five tasks. Matched controls show that the choice of negative matters more than the loss: on the same states, continued SFT recovers 21% of the gain and a random negative from the student 69%, whereas losses that keep pushing down the student's most probable mistake recover 88–97%. Its advantage over learning from the student's own samples comes mostly from states where the student never samples the expert's decision, where such methods receive no direct signal.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.