Bending the Thought: Curvature-Guided Reinforcement Learning for LLM Reasoning
Abstract
Reinforcement learning with verifiable rewards (RLVR) improves large language model reasoning using only final-answer feedback. This raises a token-level credit assignment problem: only a small subset of local computations may account for the final result, while the reward fails to indicate the locations of such computations. Existing token-level credit allocation researches, such as token entropy or update magnitude, are mostly defined on the policy side and mainly capture uncertainty or distributional change in the probability space, while neglecting the model’s internal computational processes. Inspired by the cognitive neural-manifold, we view reasoning as a trajectory in the hidden representation space. Under a smooth-trajectory idealization, we derive an upper bound on the second-order variation of a smooth scalar readout that includes a curvature-dependent term. This motivates us to treat curvature as a candidate signal for token-level update credit assignment. Building on this analysis, we introduce ARC (Adaptive Reweighting by Curvature), a curvature-guided RLVR framework that converts response-normalized curvature scores into bounded, detached weights on the token-level policy surrogate. Empirical analysis reveals a correlation between high-curvature positions and reasoning markers, alongside layer-dependent differences in entropy-based token selections. Evaluations on six mathematical reasoning benchmarks show improved performance over the baseline for both Qwen3-8B-Base and Qwen3-14B-Base, including gains of 5.98 and 8.27 percentage points on AIME 2024, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.