acceptodds
Under review as a conference paper at ICLR 2027

Sparse Where, Dense Within: Contrastive Privileged Action Distillation for Long-Horizon Agents

Abstract

Training long-horizon language model agents from terminal rewards is challenging because task outcomes provide limited guidance for individual decisions. On-policy distillation can supply dense token-level supervision using privileged training information. However, dense feedback alone does not identify which decisions warrant additional guidance, and positive guidance can conflate action preferences with the effects of added context. We introduce Contrastive Privileged Action Distillation (CPAD), which uses action comparisons supported by rollout outcomes to determine both where and how to distill. CPAD matches repeated decision states across on-policy trajectories and selects only those where downstream outcomes support a preferred action over an inferior alternative. At each selected decision, the same model re-scores its sampled response under two contexts that differ only in the guiding action. Their token-level log-probability difference augments the outcome-based advantage, making supervision selective across decisions and dense within each selected response. Privileged guidance is used only during training. Experiments on ALFWorld and WebShop with Qwen2.5 models at 1.5B, 3B, and 7B scales show competitive final performance against structured credit-assignment methods. Compared with a matched outcome-based RL baseline, CPAD improves final success and the area under the 150-update validation success curve in all six settings on ALFWorld seen tasks and WebShop. At 150 updates, final success increases by up to 24.4 and 11.4 percentage points on the two benchmarks, respectively, with 24-49% fewer training interactions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.