acceptodds
Under review as a conference paper at ICLR 2027

Turnip: Tokenizers Are Secretly Context Compressors for Transformers

Abstract

Transformer language models process text as sequences of tokens, and their self-attention cost grows quadratically with the number of tokens. As a result, tokenization largely determines the computational cost as contexts grow. However, the de facto standard tokenizer byte-pair encoding (BPE) produces token sequences longer than the information content of text requires. In this paper, we formulate tokenization as a constrained predictive-compression problem. Under this formulation, we show that the compression of BPE is limited by two constraints: its tokens are lossless and context-independent. Guided by this theory, we introduce Turnip (**t**he ne**ur**al toke**ni**zer **p**ipeline), a neural tokenizer that removes both constraints and provides direct control over the compression ratio. Turnip is model-agnostic and can be used with existing pretrained language models. On different models and benchmarks, Turnip nearly doubles the compression ratio of BPE while matching its generation quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.