Token Folding: Shortening Language-Model Sequences by Fusing Predictable Tokens into Content Words
Abstract
When using a standard tokenizer, LLM sequences carry many low-entropy tokens such as function words, punctuation, capitalization, common suffixes, etc, which we call modifier tokens whose identity in context is often low entropy, and that serve to modify the function of the ”base” word in the sentence. For example, if ”the buildings” was 3 tokens, ”[the] [building] [s]”, the main word would be [building], while [the] and [s] are modifiers. We introduce token folding: a preprocessing step that fuses such modifier tokens into an adjacent content token, so that a single sequence position now represents the ”base” word with its modifiers. The folded modifiers are predicted by a small auxiliary head, while the full transformer capacity is only spent on the ”base” word prediction. We test folded and unfolded models from 71M to 1B parameters against compute-matched no-fold twins. Folding reduces sequence length by 14.2% for a small but statistically resolved cost: +0.0064 bits per character at 356M, with a 95% interval of [+0.0056, +0.0073] over three seeds. The shorter sequence requires 14.2% fewer autoregressive forward passes, 14.2% less KV cache per unit of text, and 26.4% fewer pairwise attention interactions. At 16K context, this translates to 24.6% higher measured prefill throughput, or 17.4% after charging our current unoptimised prompt-folding implementation, with peak memory at 0.87×. The quality cost remains nearly unchanged from the compute- optimal training budget to 2.6× that budget and decreases across our 163M to 1B scale sweep. Predictable surface-form tokens can therefore be moved out of the main autoregressive sequence at a small quality cost, reducing both sequential and attention work at inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.