Revealing the Emergent Learning Mechanisms of Critic Models in PPO for LLM Reinforcement Learning
Abstract
Reinforcement learning (RL) has proven effective at improving the reasoning capabilities of large language models (LLMs) and has recently been extended to complex, long-horizon agentic tasks. This shift has renewed interest in Proximal Policy Optimization (PPO) over group-sampling-based methods such as Group Relative Policy Optimization (GRPO), owing to its finer-grained credit assignment and greater compatibility with asynchronous training. While prior work has largely focused on improving the overall performance, the learning dynamics of the critic model have received relatively little attention. In this work, we systematically investigate the emergent behaviors and learning mechanisms of the critic model in PPO across three complementary scales: *task*, *group*, and *token*. At the task level, our analysis of trajectory-level advantages reveals that the critic spontaneously develops a GRPO-like group-based normalization pattern, yet without incurring the cost of group sampling. At the group and token levels, we introduce a belief-ranking analysis to demonstrate that the critic indeed provides fine-grained learning signals, while also revealing its susceptibility to potential bias. Motivated by these findings, we identify potential failure modes of PPO and propose to decouple critic and policy training samples to better exploit informative samples. Our analysis offers a new perspective on PPO and GRPO training dynamics and mechanisms, with implications for RL algorithm design for LLMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.