StreamOE: Scaling Over Encoding with N-Gram Streams
Abstract
Sparse parameter access enables large language models (LLMs) to expand their capacity with minimal computational overhead. Over Encoding (OE) uses this principle through N-gram embedding tables, introducing an additional scaling axis beyond the backbone architecture. Specifically, OE hashes local N-grams into table indices to retrieve embeddings that augment standard token representations. An open question is how to determine appropriate sizes for these N-gram tables and allocate capacity across different N-gram orders. To address this, we conduct extensive scaling measurements using StreamOE, an architecture that directly injects N-gram information into the network's Hyper-Connections (HC). Based on these experiments and statistical corpus analysis, we identify two key findings. First, useful table capacity is influenced by dataset scale, enabling N-gram slot counts calibrated on a small proxy model to transfer to larger backbones. Second, an asymmetric capacity allocation, such as the 1:6.75 reference ratio between 2-gram and 3-gram tables, can outperform conventional uniform splits under a fixed parameter budget. We further scale StreamOE to 80B-A3B and 309B-A15B Mixture-of-Experts (MoE) models, achieving substantial improvements in both average benchmark scores and long-context performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.