Routing Under Fire? Toward Safety-Aware Multi-LLM Serving Systems
Abstract
Large language model (LLM) routers in multi-LLM serving systems balance response quality and inference cost, yet these systems also face adversarial queries that elicit harmful responses. In existing systems, routing and safety are handled separately: routers select models based on quality and cost, while defenses such as screening, translation, and regeneration are applied as fixed pipelines outside the routing decision. However, the safety and cost of a response depend jointly on the query, the model, and the defense, so a fixed defense adds unnecessary cost and quality loss on benign traffic while under-protecting against complex attacks. We therefore study safety-aware routing, which extends the routing action space from models to model–defense configurations. We introduce S²-Bench (Safety-aware Serving Benchmark), comprising 1.4 million labeled adversarial outcomes across 10 models, 16 attack settings, and 16 composite defense configurations, and show that no single defense achieves the best safety–cost trade-off across requests and models. We then propose SafeRouter, which uses risk gating to assess the risk of each query and selects a model–defense configuration accordingly, prioritizing quality for low-risk queries and the lowest-cost predicted-safe configuration for high-risk ones. On held-out attack methods, SafeRouter achieves a 0.34% attack success rate, comparable to an S²-Bench-configured Semantic Router (0.48%) at 4.5× lower adversarial cost and 2.8× lower than the strongest evaluated single-defense pipeline, while matching the benign quality of undefended IRT-Router at comparable serving cost. Code is available at https://anonymous.4open.science/r/SafeRouter-85FC.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.