acceptodds
Under review as a conference paper at ICLR 2027

Align as Act: Step-level Credit Assignment for LLM Agents via Sequential Subspace Residuals

Abstract

Reinforcement learning (RL) enables LLM agents to improve their policies by learning from environmental feedback. However, designing reliable intermediate rewards is challenging, and reliance on sparse terminal rewards can hinder efficient training. To obtain step-level signals without a separately trained evaluator, we first examine the policy's own contextual representations in trajectories on ALFWorld. We find that diverse action-feedback sequences exhibit more broadly distributed hidden-state variation than repetitive sequences. Connecting the geometry of contextual representations with classical linear innovations, we propose Align as Act (AAA), a representation-based method for constructing step-level learning signals. Processing post-action contextual representations in trajectory order, AAA projects each representation onto a subspace of directions retained from earlier interactions within the same trajectory. This separates the history-aligned component from an orthogonal innovation. The innovation energy is then transformed into an intrinsic reward signal, normalized within successful trajectories, and used to construct step advantages. We instantiate this construction as GRPO+AAA and extend it with a GiGPO-style anchor-group correction to obtain GiGPO+AAA. Experiments on ALFWorld and WebShop with Qwen2.5-1.5B-Instruct and Qwen2.5-7B-Instruct show improvements over the corresponding GRPO and GiGPO baselines. With the 1.5B policy, GRPO+AAA improves the ALFWorld success rate by 11.9 percentage points and increases the WebShop score and success rate by 8.8 points and 14.9 percentage points, respectively. Additional experiments and analyses address AAA's performance with an alternative policy backbone, the use of NNLS for failure-side credit assignment, and the role of the history basis in step-reward construction. These results suggest that trajectory-relative representation geometry provides a practical basis for step-level credit assignment in outcome-supervised LLM agent training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.