acceptodds
Under review as a conference paper at ICLR 2027

Bending the Thought: Curvature-Guided Reinforcement Learning for LLM Reasoning

Abstract

Reinforcement learning with verifiable rewards (RLVR) improves large language model reasoning using only final-answer feedback. This raises a token-level credit assignment problem: only a small subset of local computations may account for the final result, while the reward fails to indicate the locations of such computations. Existing token-level credit allocation researches, such as token entropy or update magnitude, are mostly defined on the policy side and mainly capture uncertainty or distributional change in the probability space, while neglecting the model’s internal computational processes. Inspired by the cognitive neural-manifold, we view reasoning as a trajectory in the hidden representation space. Under a smooth-trajectory idealization, we derive an upper bound on the second-order variation of a smooth scalar readout that includes a curvature-dependent term. This motivates us to treat curvature as a candidate signal for token-level update credit assignment. Building on this analysis, we introduce ARC (Adaptive Reweighting by Curvature), a curvature-guided RLVR framework that converts response-normalized curvature scores into bounded, detached weights on the token-level policy surrogate. Empirical analysis reveals a correlation between high-curvature positions and reasoning markers, alongside layer-dependent differences in entropy-based token selections. Evaluations on six mathematical reasoning benchmarks show improved performance over the baseline for both Qwen3-8B-Base and Qwen3-14B-Base, including gains of 5.98 and 8.27 percentage points on AIME 2024, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.