State Transition Matters: Self-Distilled Credit Assignment for Multi-Turn Agentic RL
Abstract
Reinforcement learning (RL) has improved large language models (LLMs) on verifiable tasks, but training multi-turn interactive agents remains challenging due to sparse outcome rewards and coarse credit assignment. Existing methods mainly rely on outcome reward back-propagation to assign credit, ignoring the rich task-relevant information hidden in state transitions, which naturally reveal the local consequences of actions. Based on this observation, we propose SDIAL, a novel Self-Distilled Advantage reshaping method for multi-turn credit assignment in reinforcement Learning. In SDIAL, the LLM of student mode interacts with the environment, while the same LLM of teacher mode leverages the state transitions to provide hindsight evaluation of the student's actions. The difference between the teacher's and student's output distributions is distilled into token-level credit signals to refine advantage estimates at each decision step. To maintain the effectiveness of self-distillation and stabilize policy optimization, SDIAL introduces an asymmetric reshaping mechanism that applies different operators to reshape positive and negative advantages. Experiments on multiple long-horizon interactive benchmarks show that SDIAL improves training stability and robustness while achieving strong task performance across diverse environments. Code is available at https://anonymous.4open.science/r/SDIAL.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.