Multi-Level Transformer: A Minimal Transformer Extension for Learned Sequence Chunking
Abstract
Byte-level sequence models waste significant capacity on the inherent compressibility of raw data. Modern tokenizers address this by chunking byte sequences into tokens with roughly uniform information density, a heuristic central to scaling large language models. However, tokenizers are static, trained once, and blind to the downstream model. The Multi-Level Transformer (MLT) introduces a minimal modification to standard Transformers, augmenting them with hierarchical processing at different temporal resolutions controlled by a target compression rate. Unlike prior dynamic-chunking approaches that rely on ratio-balancing objectives or boundary heuristics, MLT enforces capacity by construction and recovers a standard dense Transformer at full rate. Applied to byte-level language modeling, MLT outperforms prior chunking mechanisms under matched compute and matches a tokenized Transformer trained on the same data, learning compression rates that reflect the differing information densities of English, Chinese, and DNA sequences. The discovered chunk boundaries closely match those of trained tokenizers on both English and Chinese, emerging purely from the language-modeling objective. Applied to token-level language modeling, MLT learns to skip 20–30% of tokens while maintaining baseline performance, yielding significant inference savings. Finally, MLT outperforms natural language tokenizers on DNA sequence modeling, demonstrating the flexibility of learned chunking beyond language domains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.