DCGPO: Distributional Critic Guided Policy Optimization for Multi-Turn LLM Agents
Abstract
Agentic Reinforcement Learning (Agentic RL) for large language model agents faces challenges in credit assignment and exploration efficiency. Stochastic generation does not necessarily yield effective environment exploration: repeated trajectories and limited state coverage can reduce the effective use of interaction budgets. We propose Distributional Critic Guided Policy Optimization (DCGPO), an actor–critic framework that combines turn-level distributional value learning with distribution-guided exploration. The critic supports turn-level advantage estimation and provides action-conditioned return means and dispersion to guide training-time sampling. By ranking policy-generated candidate turns before execution, DCGPO directs environment interactions toward promising behaviors while allowing exploration beyond mean-greedy selection. Our fixed-policy analysis shows that categorical critic learning preserves the form and statistical order of the scalar turn-level error bound for mean-induced advantage estimation. Experiments on embodied and web-based interactive tasks across multiple model scales demonstrate the effectiveness of DCGPO, while component ablations and visitation analyses examine its effects on task performance and exploration.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.