Learning from Decision Contrasts: History-Aligned Self-Distillation for Agentic Reinforcement Learning
Abstract
Agentic on-policy self-distillation (OPSD) mitigates coarse credit assignment in outcome-based reinforcement learning (RL) by using privileged information (PI) at the trajectory or step level to guide self-teacher rescoring. However, both levels have limitations: trajectory-level supervision can weaken as student response prefixes grow, while step-level PI may be inappropriate for the target turn. Our key observation is that turns with similar interaction histories exhibit substantial contrasts in actions and outcomes, providing a natural basis for localized and informative PI construction. We propose History-Aligned Self-Distillation (HASD), a unified framework that leverages history-aligned decision contrasts to provide fine-grained and contextually appropriate OPSD supervision. HASD first groups turns by their recent observation histories and aggregates empirical action outcomes to identify those with the highest and lowest outcomes. It then constructs two PI views from these action sets with opposite high/low label assignments to produce comparative OPSD signals that mitigate the negative tendency of one-sided signals. Finally, because the comparative signal is not directly verified against task outcomes, an agreement gate retains it only when its turn-level ordering agrees with GRPO advantages. Across 1.7B, 3B, and 7B models, HASD ranks first in seven of nine settings and exceeds the strongest baselines by up to 14.1, 9.3, and 1.7 points on ALFWorld, WebShop, and Search-QA, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.