acceptodds
Under review as a conference paper at ICLR 2027

NikTok: Faster Tokenizers Need Smarter Network Hardware, Not CPUs

Abstract

Tokenization running on the host CPU competes for the same cycles as agentic workflows with applications increasingly processing longer input contexts while relying on smaller models and tokenization delays making up a larger share of end to end inference. This overhead comes primarily from the host's network stack and OS, which handles every request before it reaches the tokenizer. To eliminate this overhead and free host CPU resources, we introduce NikTok, a system that runs tokenization entirely on network hardware. NikTok accelerates the pipeline, splitting each request between many parallel hardware execution units, overlapping computation and network transfer, with tokenized output directly being written to GPU memory invoking the host CPU or the OS. Our evaluation using real network hardware across four widely used tokenizers (Qwen3-Embedding-0.6B, Harrier-270M, EmbeddingGemma-300M, and mmBERT-Embed) demonstrates that NikTok cuts average inference latency by up to 9.8× and increases throughput by up to 883%, while freeing up 83.5% of host CPU capacity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.