acceptodds
Under review as a conference paper at ICLR 2027

VTD: Distilling Symbolic Task Structure for Long-Horizon Agent Learning

Abstract

Solving long-horizon tasks requires agents to make a sequence of interdependent decisions before receiving a reliable outcome signal. On-policy distillation supplements this sparse feedback with token-level guidance from a stronger teacher on trajectories generated by the student agent. However, the teacher conditions each prediction on the student's interaction history, so earlier mistakes can distort all subsequent guidance. Our analysis further shows that even teachers that can solve a task on their own increasingly favor the failing student's actions when conditioned on this history. We therefore propose Verifiable Task Decomposition (VTD), which separates task decomposition from trajectory execution using a symbolic progress graph constructed by the teacher before interaction. During rollout, the environment turns completed graph milestones into verified progress that complements the terminal outcome in policy optimization. Graph distillation transfers this decomposition capability to the student, allowing the student to generate progress graphs without teacher access at test time. Experiments on ALFWorld, WebShop, and TextCraft show that VTD accelerates policy learning and outperforms the strongest Qwen2.5-7B baseline by 1.8, 3.4, and 8.0 points, respectively. VTD also exhibits strong structural generalization beyond the training distribution, reaching 92.8% success on held-out ALFWorld task families compared with 68.0% for the strongest baseline and extending to deeper TextCraft recipe hierarchies.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.