acceptodds
Under review as a conference paper at ICLR 2027

Prefix-LM: Prefix-Invariant Transformers for Width-Adaptive Progressive Inference

Abstract

LLMs require a significant amount of resources during autoregressive generation, but not all tokens are equal. Token prediction is a task of varying difficulty, but monolithic architectures allocate the same compute regardless of the difficulty of prediction. Progressive inference, where representations are incrementally computed and refined as needed, can address this challenge. This has traditionally been achieved through early-exiting methods that operate along the depth dimension of the network. Elastic networks that operate along the width dimension can select a different computational budget per input, but the resulting outputs cannot be progressively refined. We identify the missing property as prefix invariance: the intermediate representations of smaller submodels are prefixes of larger ones. To this end, we propose Prefix-LM: an architectural design that achieves prefix invariance by enforcing block-triangular maps, enabling KV cache sharing and full reuse of intermediate computations. We apply Prefix-LM at the 1B scale, demonstrating practical speedups and establishing prefix invariance as a useful foundation for adaptive computation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.