acceptodds
Under review as a conference paper at ICLR 2027

Hierarchical Self-Distillation for Agentic Reinforcement Learning

Abstract

On-policy self-distillation (OPSD) has emerged as a promising post-training approach that complements sparse outcome rewards in agentic reinforcement learning with dense guidance from a privileged self-teacher. Nevertheless, we reveal that the utility of teacher supervision in multi-turn OPSD is hierarchically non-uniform: step-level supervision utility is influenced by both accumulated state drift and future teacher–student divergence, while at the token level, under a limited support budget, a fixed-size support cannot maintain consistent fidelity to full-distribution supervision as teacher–student discrepancy varies. To address this, we propose HiSD, a hierarchical self-distillation framework that combines progress-aware step supervision with support-adaptive token supervision, allocating step-level distillation according to relative divergence progress and adapting the distributional scope of token-level supervision based on teacher–student discrepancy. Across ALFWorld, Search-QA, and WebShop with Qwen2.5-Instruct models, HiSD consistently outperforms GRPO and recent OPSD-based baselines. Notably, on tasks with strong dependencies across interaction steps, HiSD outperforms SDAR by an average of points on Pick2 and points on multi-hop QA, while achieving higher teacher budget utilization on WebShop. Code available: https://anonymous.4open.science/r/HiSD.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.