acceptodds
Under review as a conference paper at ICLR 2027

A Probe Critic Matches a Policy-Sized Critic in Agentic PPO

Abstract

Proximal policy optimization (PPO) for language models typically uses a separate value network comparable in size to the policy, and agentic reinforcement learning has kept that design for trajectories that run to thousands of tokens across many turns. These policy-sized critics add considerable memory and compute and are pretrained before use. We show that a small probe trained on a language model agent's own activations can serve as its critic, at four orders of magnitude fewer parameters. We demonstrate that training the policy with such a probe nearly matches the performance of a critic the size of the policy, and holds that performance when the trajectories it learns from are several updates off-policy. We further find probe critics robust to how they are fit. A cold-start probe learns value alongside the policy, and a warm-start probe that never co-learns can still train the policy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.