XSR: Rethinking Where Semantic Routing Happens
Abstract
Semantic routing is an essential component of large-scale, multi large language model (LLM) serving, enabling requests to be dynamically directed to the right models based on their content. However, existing semantic routers operate as application-layer intermediaries, requiring requests to incur communication overhead before being forwarded to selected models. This design introduces severe overhead from application-layer processing and communication with external router processors. We present eXpress Semantic Router (XSR), a semantic router architecture that achieves high-throughput and low-latency by running latency-sensitive routing components, such as request processing, signal extraction, and decision policy, in the network layer. Our experiments show that across different routing functions (character (n)-gram, BM25, and distilled intent routing), XSR delivers up to 37.2 higher throughput across the evaluated concurrency ranges while reducing average latency by up to 79.7% compared to a state-of-the-art baselines, such as vLLM Semantic Router.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.