acceptodds
Under review as a conference paper at ICLR 2027

DL-OPSD: Dual-Level Selective On-Policy Self-Distillation for Skill Internalization

Abstract

On-policy self-distillation combined with reinforcement learning offers a route to internalizing external skills in language agents, removing the need for skill retrieval at inference. However, local teacher–student prediction gaps alone are insufficient to determine which behaviors deserve reinforcement through distillation. We find that, when teacher–student gaps are comparable, original actions yield positive average gains in continuation success over student-sampled alternatives on positive-advantage trajectories, but negative gains on negative-advantage trajectories. This observation suggests that trajectory outcomes inform supervision selection beyond local prediction gaps, motivating Dual-Level Selective On-Policy Self-Distillation (DL-OPSD). At the trajectory level, DL-OPSD applies auxiliary distillation only to trajectories with positive group-relative advantage, while retaining reinforcement learning updates across all trajectories. At the token level, it selects positions where the teacher assigns a higher probability to the sampled token than the student does, and weights the supervision according to student uncertainty and teacher certainty. Skills are provided only to the teacher during training. Across ALFWorld, Search-QA, and WebShop, DL-OPSD outperforms GRPO and SDAR on the primary task metrics, improving Qwen3-1.7B's ALFWorld success rate from 60.9% under GRPO to 70.3%. Controlled experiments show that trajectory selection improves gradient alignment between distillation and reinforcement learning, while token-level weighting further improves task performance, supporting the complementary roles of the two levels.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.