Learning What the Group Cannot Tell: Decomposed Critics for LLM Reinforcement Learning
Abstract
Reinforcement learning (RL) has become an important approach for post-training large language models (LLMs). In many LLM RL tasks, feedback is sparse or available only at the end of a trajectory, making credit assignment across intermediate states important for effective policy optimization. Proximal Policy Optimization (PPO) addresses this problem with a learned value function, which estimates the expected return of intermediate states and provides state-dependent advantages for policy updates. However, modern LLM-based RL algorithms typically collect multiple rollouts for each prompt, thereby exposing group-level information that conventional critics do not explicitly exploit. In this work, we revisit critic learning in such grouped settings.Instead of learning the full state value directly from individual terminal outcomes, We decompose value estimation into two complementary components: a prompt-level success prior estimated from grouped rollouts and history-dependent evidence learned by the critic. Based on this decomposition, we introduce **G**roup-**A**nchored **V**alue **E**stimation (GAVE), which trains the critic with a group-conditional objective and reconstructs standard state values by combining the learned correction with an empirical prompt prior. We further prove that estimation errors in the empirical prior are not amplified in expectation during reconstruction, yielding a finite-sample error bound. Experiments on LLM agent tasks show that GAVE consistently improves both overall and within-prompt value estimation compared with conventional critics and information-matched baselines, while also leading to improved policy optimization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.