TreePPO: Learning from Trees and Exploring with Critics for Multi-Turn LLM Agents
Abstract
Multi-turn LLM agents typically make sequential decisions over long interaction trajectories, with reward signals often available only upon task completion. Existing reinforcement learning (RL) methods largely rely on these sparse terminal rewards, making it difficult to distinguish the contributions of intermediate actions to final task success. Although a critic can mitigate this limitation by estimating state values learning, accurate values for intermediate states remains challenging when the critic is trained primarily on sparse terminal rewards. To address this challenge, we propose TreePPO, a tree-search-based actor-critic RL framework that establishes a bidirectional feedback loop between rollout exploration and critic learning for continual improvement. TreePPO aggregates rollouts with shared action prefixes to provide fine-grained value supervision, while the learned critic guides adaptive rollout allocation toward tasks with greater exploration potential. Specifically, TreePPO organizes trajectories sharing ordered action prefixes into an Action-Prefix Value Tree (APVT). By aggregating terminal rewards from different continuations of each shared prefix, the APVT provides empirical value targets for intermediate states, enabling finer-grained critic learning. APVT provides supervision for critic warm-up, while subsequent on-policy rollouts continually expand the state space covered by APVT. We further use the critic's estimates of task difficulty to adaptively allocate the rollout budget, assigning more samples to moderately difficult tasks that offer greater exploration value. Extensive experiments across Multi-hop QA, ALFWorld, and WebShop show that TreePPO achieves the best overall performance among the evaluated methods with both Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct actors. With Qwen2.5-7B-Instruct, TreePPO achieves an average exact match (EM) of 36.60% across four Multi-hop QA benchmarks and an ALFWorld success rate of 86.07%, exceeding the strongest baselines in these two settings by 2.15 and 5.97 percentage points, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.