Consumption Is Unnecessary: Learning a Value Estimate Is Enough for Group-Relative Agent RL
Abstract
A common pattern in training LLM agents is to let the policy estimate its own progress or its chance of success and then use that estimate — as an advantage, a reward, or a rule for choosing actions. Whether the gain comes from learning the estimate or from using it has not been separated. We separate the two. Our method, RAISE, trains the estimate with a listwise rank loss whose only label is the trajectory outcome, and then never reads it: not in the reward, not in the advantage, not at inference. The loss reaches the policy only through the shared parameters. On ALFWorld, Sokoban and WebShop, with three training seeds each, this objective alone beats the published GiGPO baselines by , and points at the same budget, and reaches their final success within – of it. Where we test consumption — a forward difference of the estimate injected into the advantage at two doses, and three rules for choosing the action with it at inference — it adds nothing we can detect: the injection is flat at both doses and the selection rules do not help. Randomising the targets removes the gain, so what matters is that the targets are correct. What the loss teaches is which states are promising, not which action to take: inside a state, the order it gives the candidate actions is no better than one constant per state. The policy's hidden state already carries the outcome before the loss is added. The practical result is a recipe that needs no demonstrations and no extra forward pass.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.