From Tokens to Contexts: Progressive N-Gram Hash Routing for Sparse MoE Language Models
Abstract
Early-layer hash routing has emerged as a practical alternative to fully learned expert selection in sparse Mixture-of-Experts (MoE) pretraining, but existing token-ID hashing answers only where deterministic routing can be used, not what routing key should be used within the shallow block. Because every occurrence of the same token is mapped to the same expert set, token-ID hashing is context-blind, can concentrate frequent-token traffic, and ignores the progressive change in representation granularity with depth. We propose ProGram, a layer-aware multi-prime n-gram hash routing method that treats shallow MoE routing as depth-aware routing-key design. ProGram replaces token-ID routing atoms with ordered local n-gram contexts, constructs K distinct experts through multi-prime hashing, stable deduplication, and fallback probing, and applies a progressive routing-order schedule ν = (1, 2, 3) across the first three MoE layers. This yields a gradual transition from lexical identity to short-range contextual patterns: hashing determines expert identities in shallow layers, learned router scores compute adaptive mixture weights over the selected experts, and deeper layers retain standard learned top-K routing. We pretrain a 5B-parameter MoE model with 412M activated parameters per token on approximately 100B tokens and evaluate it on 12 knowledge, reasoning, and code benchmarks. In controlled comparisons, ProGram₁₂₃ obtains the best result on 10 of the 12 benchmarks relative to dense-replacement baselines and achieves higher scores than a fully learned-routing MoE baseline on 11 benchmarks, with a tie on MGSM. It also yields lower language-modeling loss and lower sequence-level load-balancing loss during pretraining. These results support progressively contextualized hash routing as a useful shallow-layer routing prior. Code is available at https://anonymous.4open.science/r/Program-FE0E/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.