acceptodds
Under review as a conference paper at ICLR 2027

UMCTS-GRPO: Uncertainty-Aware Monte Carlo Tree Search for LLM Agent Reinforcement Learning

Abstract

Tool-using language agents make sequences of decisions but are typically rewarded only for the final outcome. Tree-structured rollouts expose alternative decisions at shared histories, turning terminal rewards into step-level credit, yet existing tree-based agent RL expands branches at random and trusts every comparison equally. We introduce UMCTS-GRPO, which uses a single backed-up branch uncertainty to decide both where to collect rollout evidence and how much to trust it. During collection, uncertainty-aware Monte Carlo tree search steers simulations toward unresolved branches; during optimization, the same statistic weights sibling comparisons through an uncertainty-weighted baseline and a confidence gate, yielding step-level advantages without a critic, process reward model, or step-level annotation. In short: explore what you doubt; learn from what you have resolved. Experiments on seven retrieval-based QA benchmarks and the interactive ALFWorld environment, across five backbones from the Qwen2.5 and Llama3.2 families, show that UMCTS-GRPO consistently improves over strong agent RL baselines and remains effective under tight rollout budgets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.