BridgeSTOP: Bridging Streaming Tokenization and LLM Prefill
Abstract
Streaming computation has been increasingly adopted across LLM serving to improve efficiency. However, existing streaming methods mainly optimize individual serving stages, while the tokenization-to-prefill path remains limited by batch-bounded inter-request execution, insufficient bridging between the two stages, and transient efficiency mismatch across them. We propose , a framework that bridges streaming tokenization and LLM prefill to improve, preserve, and propagate streaming tokenization efficiency downstream and ultimately translate it into end-to-end serving gains. first introduces batch-free asynchronous parallel tokenization, which extends streaming tokenization from intra-request processing to inter-request execution to reduce batch-bounded synchronization waiting and improve tokenization efficiency. We then establish a token pool that continuously collects streaming token outputs and controls when and how many available tokens are released to downstream prefill, preserving and propagating streaming tokenization efficiency. When upstream tokenization temporarily falls behind, we introduce speculative tokenization as a fast replenishment mechanism to replenish transient token deficits and prevent downstream prefill starvation, while the longest-match speculation strategy improves speculation accuracy and reduces rollback overhead. Experiments on Qwen3-8B across five diverse workloads show that consistently improves sustainable serving capacity and reduces latency, effectively translating streaming tokenization efficiency into end-to-end serving gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.