acceptodds
Under review as a conference paper at ICLR 2027

Distribution-Preserving Response Length Separation for Language Models

Abstract

Response length is a major source of uncertainty in serving large language models. Under stochastic decoding, the same request can produce responses that differ several-fold in length: for prompts from a single dataset, within-prompt variance in response length can exceed between-prompt variance. This uncertainty undermines efficient resource management. Predicting length from the prompt cannot address within-prompt variation, while steering generation towards a target length changes the model's output distribution. We introduce **Length-Separated Sampling**, which divides generation into short and long regimes chosen at random before generation. We construct both regimes from the frozen target model so that their mixture is exactly the original distribution. The regime is therefore known before generation and informative about response length, while the responses users receive remain distributed exactly as under ordinary sampling. The long regime reliably generates longer responses than the short regime, with mean response length up to 2.8 times that of the short regime.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.