acceptodds
Under review as a conference paper at ICLR 2027

Disentangled Token-Level Credit Assignment for Efficient Reasoning

Abstract

Large reasoning models (LRMs) rely on explicit multi-step reasoning to substantially improve complex problem solving but are also prone to overthinking, which consumes computation with no gain in accuracy and can even derail otherwise correct trajectories. Existing approaches primarily mitigate overthinking through length penalties. However, such sequence-level signals penalize an entire long response uniformly, which suppresses useful reasoning along with redundant computation. To address this limitation, we propose Per-Objective Asymmetric Credit Estimation (PACE), which disentangles correctness from efficiency and assigns their learning signals separately at the token level, charging each generation decision for the computation it induces while reinforcing correct reasoning. Specifically, PACE employs intra-prompt correctness credit to characterize relative solution quality among responses to the same prompt, while inter-prompt rate credit apportions success gains and remaining-compute penalties to individual tokens. We further propose an asymmetric correctness budget that bounds negative efficiency credit on correct trajectories and retains the full remaining-compute penalty on erroneous trajectories. Theoretically, we establish that the inter-prompt rate credit constitutes a token-level cross-fitted estimator of the pooled solve-throughput gradient, yielding a principled decomposition of the global accuracy-compute objective into token-wise success gains and remaining-compute penalties. Extensive experiments across ten reasoning benchmarks and two model scales demonstrate that PACE reduces generation length up to 64% while improving average accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.