acceptodds
Under review as a conference paper at ICLR 2027

Fixed Length Penalties: Accuracy–Length Trade-offs in Mathematical Reasoning

Abstract

Length-aware reinforcement learning can shorten reasoning, but assessing adaptive rewards requires a strong account of what a simple fixed cost already achieves. We study a constant penalty on capped, normalized completion length in reinforcement learning with verifiable rewards. Across eight paired training seeds of DeepSeek-R1-Distill-Qwen-1.5B, a coefficient of 0.3 reduces mean generated tokens by 26.9%, 12.0%, 2.8%, and 7.8% on GSM8K, MATH-500, AIME, and AMC, with mean accuracy changes of −0.80, +0.63, +0.21, and +2.11 percentage points. Every seed pair yields shorter generations on every benchmark, and tokens per correct response fall by 26.2%, 12.7%, 3.8%, and 11.1%. Without problem-specific budgets, difficulty estimates, or online penalty adaptation, the fixed cost captures 84.9–97.3% of the largest mean token saving achieved by the tested adaptive controls. The largest absolute savings occur in intermediate difficulty groups, while difficult AIME problems dominate the remaining generation volume. These results establish a simple and competitive accuracy–efficiency reference for evaluating adaptive length control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.