Every Second Counts: Quality-Cost-Aware LLM Routing for Low-Latency Concurrent Multi-LLM Serving
Abstract
As the Large Language Model (LLM) ecosystem grows, LLM routing is emerging as an important approach to balancing response quality and inference cost. Since real-world LLM serving involves continuously arriving requests, backlog-induced latency and quality degradation can no longer be overlooked. However, existing methods optimize only quality and cost, largely because they omit backlog-induced latency. In addition, they directly select a model without an admission mechanism, leading to quality degradation. To address these limitations, we propose RTRouter, a quality-cost-aware LLM routing framework that couples workload-awareness with quality admission. Specifically, we represent LLMs' unfinished work backlog as workload and update it online to avoid missing backlog-induced latency. Furthermore, we adaptively adjust the degree of pruning based on the quality vector and prune candidate models to obtain the feasible set, avoiding quality degradation. Extensive experiments on comprehensive routing benchmarks demonstrate that RTRouter improves the quality-cost-latency trade-off and effectively eliminates quality degradation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.