acceptodds
Under review as a conference paper at ICLR 2027

TAAD: Teachability-Aware Agentic Distillation for Multi-Step Credit Assignment

Abstract

Multi-step agentic reinforcement learning trains agents to complete long-horizon trajectories under a sparse terminal return, where Group Relative Policy Optimization (GRPO) broadcasts one sequence-level advantage uniformly to every token. Token-level teacher-student self-distillation sharpens this signal, yet single-token probability gaps can conflate distributional disagreement with incidental fluctuation and do not exploit step structure. In this paper, we propose Teachability-Aware Agentic Distillation (TAAD), a bounded two-level advantage reweighting method. TAAD replaces the single-token gap with a teacher-mass-weighted disagreement signal over the shared teacher-student top- support, aggregates this signal over action spans, and gates its rising inter-step transitions by student action surprisal to form a step-priority score. A maximum-entropy softmax step budget is then refined into per-token weights. The reshaped advantage preserves the sign of the sequence-level advantage at every valid token while allocating different credit magnitudes across a trajectory. Across both Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct backbones, TAAD consistently outperforms all baselines on the three multi-step benchmarks of ALFWorld, Search-QA, and WebShop.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.