acceptodds
Under review as a conference paper at ICLR 2027

TreePPO: Learning from Trees and Exploring with Critics for Multi-Turn LLM Agents

Abstract

Multi-turn LLM agents typically make sequential decisions over long interaction trajectories, with reward signals often available only upon task completion. Existing reinforcement learning (RL) methods largely rely on these sparse terminal rewards, making it difficult to distinguish the contributions of intermediate actions to final task success. Although a critic can mitigate this limitation by estimating state values learning, accurate values for intermediate states remains challenging when the critic is trained primarily on sparse terminal rewards. To address this challenge, we propose TreePPO, a tree-search-based actor-critic RL framework that establishes a bidirectional feedback loop between rollout exploration and critic learning for continual improvement. TreePPO aggregates rollouts with shared action prefixes to provide fine-grained value supervision, while the learned critic guides adaptive rollout allocation toward tasks with greater exploration potential. Specifically, TreePPO organizes trajectories sharing ordered action prefixes into an Action-Prefix Value Tree (APVT). By aggregating terminal rewards from different continuations of each shared prefix, the APVT provides empirical value targets for intermediate states, enabling finer-grained critic learning. APVT provides supervision for critic warm-up, while subsequent on-policy rollouts continually expand the state space covered by APVT. We further use the critic's estimates of task difficulty to adaptively allocate the rollout budget, assigning more samples to moderately difficult tasks that offer greater exploration value. Extensive experiments across Multi-hop QA, ALFWorld, and WebShop show that TreePPO achieves the best overall performance among the evaluated methods with both Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct actors. With Qwen2.5-7B-Instruct, TreePPO achieves an average exact match (EM) of 36.60% across four Multi-hop QA benchmarks and an ALFWorld success rate of 86.07%, exceeding the strongest baselines in these two settings by 2.15 and 5.97 percentage points, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.