acceptodds
Under review as a conference paper at ICLR 2027

DCGPO: Distributional Critic Guided Policy Optimization for Multi-Turn LLM Agents

Abstract

Agentic Reinforcement Learning (Agentic RL) for large language model agents faces challenges in credit assignment and exploration efficiency. Stochastic generation does not necessarily yield effective environment exploration: repeated trajectories and limited state coverage can reduce the effective use of interaction budgets. We propose Distributional Critic Guided Policy Optimization (DCGPO), an actor–critic framework that combines turn-level distributional value learning with distribution-guided exploration. The critic supports turn-level advantage estimation and provides action-conditioned return means and dispersion to guide training-time sampling. By ranking policy-generated candidate turns before execution, DCGPO directs environment interactions toward promising behaviors while allowing exploration beyond mean-greedy selection. Our fixed-policy analysis shows that categorical critic learning preserves the form and statistical order of the scalar turn-level error bound for mean-induced advantage estimation. Experiments on embodied and web-based interactive tasks across multiple model scales demonstrate the effectiveness of DCGPO, while component ablations and visitation analyses examine its effects on task performance and exploration.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.