acceptodds
Under review as a conference paper at ICLR 2027

GRAR: Graph-Rubric-guided Agentic Reward for Reinforcement Learning of LLM Agents

Abstract

Reinforcement learning for large language model (LLM) agents is naturally super-vised by the environment’s terminal signal: whether the task was completed. Under group-relative policy optimization (GRPO), terminal supervision is uninformative in two distinct regimes. In an all-failed group, every sampled trajectory receives the same reward, so the group produces no policy-gradient signal. In a mixed group containing both successful and failed trajectories, the terminal reward distinguishes their outcomes but does not identify which decisions contributed to success or failure. We introduce GRAR (Graph-Rubric-guided Agentic Reward), a post-hoc credit-assignment framework that preserves the terminal reward while providing process-level feedback in both regimes. For groups containing at least one success, graph settlement assigns step credit from verified outcomes represented in a group-local semantic state graph. For all-failed groups, rubric settlement scores a small set of graph-selected steps with a single guarded LLM-judge call. Trajectory advantages remain determined solely by terminal rewards, while process credit enters as a bounded step-aligned bonus; under the deployed configuration, this preserves the ordering between successful and failed trajectories. A single environment-agnostic adapter spans τ 2-bench, WebShop, and ALFWorld. Across the evaluated bench-mark settings, the best GRAR variant improves on the terminal-only baseline on both metrics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.