Fixed Length Penalties: Accuracy–Length Trade-offs in Mathematical Reasoning
Abstract
Length-aware reinforcement learning can shorten reasoning, but assessing adaptive rewards requires a strong account of what a simple fixed cost already achieves. We study a constant penalty on capped, normalized completion length in reinforcement learning with verifiable rewards. Across eight paired training seeds of DeepSeek-R1-Distill-Qwen-1.5B, a coefficient of 0.3 reduces mean generated tokens by 26.9%, 12.0%, 2.8%, and 7.8% on GSM8K, MATH-500, AIME, and AMC, with mean accuracy changes of −0.80, +0.63, +0.21, and +2.11 percentage points. Every seed pair yields shorter generations on every benchmark, and tokens per correct response fall by 26.2%, 12.7%, 3.8%, and 11.1%. Without problem-specific budgets, difficulty estimates, or online penalty adaptation, the fixed cost captures 84.9–97.3% of the largest mean token saving achieved by the tested adaptive controls. The largest absolute savings occur in intermediate difficulty groups, while difficult AIME problems dominate the remaining generation volume. These results establish a simple and competitive accuracy–efficiency reference for evaluating adaptive length control.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.