Tree-Aware Posterior Distillation for Long-Horizon Agentic Reinforcement Learning
Abstract
Reinforcement learning with verifiable rewards has become a scalable post-training paradigm for large language models. However, training agents over long-horizon interactions presents two challenges: using only terminal rewards makes turn-level credit assignment sparse and high-variance, while repeated updates on a fixed rollout batch introduce policy staleness as the current policy drifts from the behavior policy. Previous work has used learned critics, process reward models, and on-policy distillation to provide denser credit, but these methods need a value model, process labels, or a stronger external teacher, extra components that increase training complexity. Importance sampling and clipping restrict policy drift, but they do not specify a finite target that stops sequential updates from overshooting. In this work, we introduce TreeOPSD, a critic-free, tree-aware posterior self-distillation framework for long-horizon agentic reinforcement learning. By treating each interaction turn as a distinct tree node, TreeOPSD branches alternative continuations from visited turn-level prefixes and recursively backs up their terminal rewards into implicit turn-level credit. It then converts each credit into a closed-form, bounded likelihood-ratio teacher derived from outcome rewards relative to the frozen behavior policy. Matching the student policy to this teacher mitigates policy drift under sequential mini-batch optimization with a finite target that stops sequential updates from overshooting. This matching recovers the policy gradient direction at the behavior policy. Empirically, TreeOPSD improves agentic capabilities across several competitive benchmarks and model scales, especially for long-horizon tasks, e.g., raising the average accuracy of Qwen3-30B-A3B-Thinking-2507 by 4.97 points over GRPO on the Deep Research setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.