acceptodds
Under review as a conference paper at ICLR 2027

When Return Bounds Do Not Control Update Weights: Prediction Rewards in Multi-Turn LLM Agents

Abstract

Small coefficients and episodic return bounds are commonly used to limit the influence of auxiliary rewards. We show that neither necessarily controls the weights used by a normalized policy update. We study a next-observation prediction reward defined by differences between consecutive rule-based prediction scores. Its contribution telescopes across an episode and is bounded in magnitude by a coefficient λ. Nevertheless, adding this reward on ALFWorld produces a striking prediction–task divergence: across Qwen3-1.7B, 4B, and 8B, the rewarded prediction score approaches one while task success approaches zero. We show how the return-level guarantee can fail in the studied step-pooled GRPO implementation. Interaction steps from multiple trajectories are standardized before the prediction increments telescope. When task returns tie or have little spread, residual prediction increments can rank the samples; if their variation controls the normalization denominator, within-group standard-deviation division approximately cancels λ, yielding order-one advantage weights. Empirically, a tenfold coefficient sweep does not prevent collapse. In a matched Qwen3-4B comparison, removing only within-group standard-deviation division restores baseline-range learning, while filtering all-failure groups provides a near-baseline but distribution-changing rescue. The failure also transfers to WebShop, where the mean-only intervention restores performance. These results show that auxiliary rewards should be evaluated after estimator transformation: return-level bounds alone do not guarantee update-level safety.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.