Length Is Not Just Bias: Marginal Token Utility in LLM Post-Training
Abstract
Response length is often treated as a source of bias in LLM post-training, motivating methods that globally suppress or normalize its influence on learning. We study this issue in preference-based post-training, where our analysis of training data shows that the relationship between response length and quality may not be universal: additional length is associated with better outcomes in some contexts, contributes little in others, and can be detrimental elsewhere. To capture this context-dependent relationship, we introduce marginal token utility, which measures the gain associated with additional response tokens within each context. This perspective suggests that length should not be uniformly penalized or removed from optimization. Instead, its influence should reflect the marginal gains associated with additional tokens. Building on this insight, we propose a marginal-utility-guided objective that modulates length-related learning according to marginal token utility. Experiments across language models and alignment benchmarks show that our approach consistently outperforms existing length-debiasing methods while avoiding unnecessary verbosity. Our results suggest that response length should be viewed not merely as an optimization bias, but as a context-dependent source of marginal utility in LLM post-training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.