acceptodds
Under review as a conference paper at ICLR 2027

Token Savings Can Mask Recoverable Taskwise Regressions in Length-Controlled RL

Abstract

Length-controlled reinforcement learning (RL) aims to learn token-efficient chains of thought while preserving task accuracy. However, average length and aggregate accuracy can conceal taskwise regressions, leaving practitioners uncertain whether to allocate more inference tokens or modify training. We introduce a taskwise intervention framework that connects regression diagnosis to training decisions. It tracks the same tasks through fixed-weight inference interventions and training continuations that retain, relax, or remove length penalties, comparing recovery at matched inference cost. Across the Qwen3.5-4B logic and code experiments, strong fixed soft-overlong penalties on all responses achieve at least 30.82% token savings at equivalent aggregate accuracy, yet at least 11.0% of tasks remain degraded across the tested inference controls. In each training seed, a single modified-penalty branch recovers at least a 6.0% share of the full task panel from these persistent deficits, outperforming continued strong-penalty training at matched inference cost. Across five seeds per domain, taskwise evidence selects Relax over Keep in all ten runs, reducing prespecified held-out decision loss by 11.30% on logic and 13.77% on code on average, with 15.2% and 18.0% more inference tokens. These gains include minimum recorded monitoring and exploratory-training costs; token-heavy weights reverse this preference. The framework connects taskwise regression and recovery to continuation choices beyond aggregate token savings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.