acceptodds
Under review as a conference paper at ICLR 2027

GAIT: Geometry-Based Credit Assignment from Hidden-State Trajectories for RLVR

Abstract

Reinforcement learning with verifiable rewards (RLVR) has become a popular approach for improving LLM reasoning. However, credit assignment in mainstream RLVR methods remains coarse-grained: a sequence-level reward is typically propagated across the entire response, regardless of which reasoning steps are essential for the final outcome. Existing approaches to step-level credit assignment often rely on additional supervision that is costly to obtain, such as process reward models, or computationally expensive Monte Carlo sampling that branches from intermediate steps. In this work, we propose **G**eometry-based Credit **A**ssignment from H**I**dden-state **T**rajectories (GAIT), a plug-and-play method that derives step-level credits directly from the model's hidden-state trajectories, without requiring any external step-level supervisor. To compare the geometry of responses with different lengths and reasoning paces, GAIT aligns each response with successful responses from the same rollout group using dynamic time warping, an algorithm for matching temporal sequences that may differ in length. It then measures step-wise divergence in a low-dimensional representation space, and uses temporal changes in this divergence to reweight the sequence-level advantage across reasoning steps. Intuitively, this assigns stronger negative updates to steps in incorrect responses that increasingly diverge from successful trajectories, and stronger positive updates to atypical but successful transitions in correct responses. Across model scales and families, GAIT outperforms existing RLVR baselines with and without fine-grained credit assignment. On challenging AIME 2024-2026 benchmarks with Qwen3-8b-Bases as the base model, GAIT improves over DAPO by 2.38 points and reaches DAPO's peak performance with 1.31 fewer GPU hours, while adding negligible computational overhead to each RL update.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.