Dynamic LLM Decoding Length Prediction: Learning from Branching Futures
Abstract
Length prediction during large language model (LLM) decoding provides an important signal for allocating limited computational resources across concurrent requests. Existing decode-time predictors typically learn from sampled trajectories, each of which reveals only one realized future at every visited prefix. However, a prefix admits multiple possible futures, and its remaining-length distribution evolves as more response content is observed. The challenge is therefore to obtain informative conditional supervision without exhaustively exploring this branching future space. In this paper, we propose *BRANching-supervised Distribution prediction* (BranD), an approach for learning prefix-conditioned remaining-length distributions under a limited continuation budget. BranD combines current and historical target-model representations to update its predictions as decoding proceeds. For training, it samples independent continuations from selected prefixes to refine conditional distribution targets, while reusing their paths to expand prefix coverage. An adaptive allocation strategy combines estimated conditional variability with predictor sensitivity to determine how many continuations to sample at each prefix. We provide a fixed-feature risk analysis that connects conditional sampling error to shared prediction error and justifies the allocation strategy. Experiments across multiple models and tasks validate the effectiveness of our proposed BranD.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.