NikTok: Faster Tokenizers Need Smarter Network Hardware, Not CPUs
Abstract
Tokenization running on the host CPU competes for the same cycles as agentic workflows with applications increasingly processing longer input contexts while relying on smaller models and tokenization delays making up a larger share of end to end inference. This overhead comes primarily from the host's network stack and OS, which handles every request before it reaches the tokenizer. To eliminate this overhead and free host CPU resources, we introduce NikTok, a system that runs tokenization entirely on network hardware. NikTok accelerates the pipeline, splitting each request between many parallel hardware execution units, overlapping computation and network transfer, with tokenized output directly being written to GPU memory invoking the host CPU or the OS. Our evaluation using real network hardware across four widely used tokenizers (Qwen3-Embedding-0.6B, Harrier-270M, EmbeddingGemma-300M, and mmBERT-Embed) demonstrates that NikTok cuts average inference latency by up to 9.8× and increases throughput by up to 883%, while freeing up 83.5% of host CPU capacity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.