acceptodds
Under review as a conference paper at ICLR 2027

Bounded Exponential Reward Shaping for Adaptive Reasoning Efficiency

Abstract

Reasoning models often generate lengthy responses even when additional computation offers limited benefit. Uniform length penalties encourage conciseness but do not account for differences in problem difficulty. We study DASE (Difficulty-Adaptive Shaping for Efficiency), a reinforcement learning approach that uses the current rollout group's solve rate to adapt a capped exponential penalty on reasoning length. The penalty encourages stronger compression on reliably solved prompts and relaxes it on lower-success prompts, without requiring a separate difficulty predictor or an inference-time controller. Experiments with Qwen3.5-122B-A10B across science, mathematics, and coding benchmarks show an approximately 38.4% reduction in average response length relative to an unpenalized reference and the highest mean Overthinking-Adjusted Accuracy area among the compared methods. Ablations examine the influence of the shaping parameters, while trajectory analysis reveals longer reasoning on lower-success groups. These findings support bounded, difficulty-adaptive reward shaping as a practical approach to improving reasoning efficiency.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.