Turn-Level Critic Guidance for Training and Inference in Terminal Agents
Abstract
Terminal agents solve tasks by reasoning, executing commands, and adapting to feedback from the environment. They can make mistakes and recover within a single rollout, yet Group Relative Policy Optimization (GRPO) assigns the same advantage to every action in that rollout. Finer-grained methods typically derive advantages from a learned value function, or rely on matched or restorable states, known answers, or judges. We introduce GRPO with Turn-Level Critic Guidance (TCG-GRPO), which uses a co-trained critic to guide policy learning and action selection. The critic predicts eventual task success from the interaction history, and the change in its score across each turn provides turn-level credit that augments the GRPO advantage. We also adopt an auxiliary objective that trains the policy to predict environment observations. The resulting Qwen3-8B policy achieves 7.3% pass@1 on Terminal-Bench 2.1 and 72.8% on held-out Endless Terminals tasks, compared with 2.9% and 54.3% for GRPO. During inference, we reuse the co-trained critic for turn-level best-of-N action selection without additional training. With four candidates per turn, this achieves 79.2% pass@1 on Endless Terminals, compared with 71.6% for standard best-of-N with log-probability ranking.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.