acceptodds
Under review as a conference paper at ICLR 2027

Every Second Counts: Quality-Cost-Aware LLM Routing for Low-Latency Concurrent Multi-LLM Serving

Abstract

As the Large Language Model (LLM) ecosystem grows, LLM routing is emerging as an important approach to balancing response quality and inference cost. Since real-world LLM serving involves continuously arriving requests, backlog-induced latency and quality degradation can no longer be overlooked. However, existing methods optimize only quality and cost, largely because they omit backlog-induced latency. In addition, they directly select a model without an admission mechanism, leading to quality degradation. To address these limitations, we propose RTRouter, a quality-cost-aware LLM routing framework that couples workload-awareness with quality admission. Specifically, we represent LLMs' unfinished work backlog as workload and update it online to avoid missing backlog-induced latency. Furthermore, we adaptively adjust the degree of pruning based on the quality vector and prune candidate models to obtain the feasible set, avoiding quality degradation. Extensive experiments on comprehensive routing benchmarks demonstrate that RTRouter improves the quality-cost-latency trade-off and effectively eliminates quality degradation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.