Error-Aware Reverse Auction Mechanism for Large Language Model Routing
Abstract
Routing each query to a cost-effective large language model (LLM) is critical for balancing quality and cost, yet most routers rely on a centralized task center to predict model performance, creating an information-risk mismatch and a scalability bottleneck as the model pool grows. We formulate LLM routing as a market-based allocation problem among strategic providers and propose a routing paradigm that shifts ex-ante prediction to LLM providers via a reverse auction, where providers submit self-predicted acceptance probabilities and execution costs. To account for noisy provider predictions and center evaluations, we introduce the Error-Aware Reverse Auction Mechanism (EA-RAM), which explicitly models this Dual Error. We prove that, under a private-evaluation-belief structure, truthful effective-surplus reporting is incentive compatible in the reduced-form score space and individually rational under sellers' subjective beliefs, establish sufficient conditions for center rationality, and derive an explicit social-welfare loss bound. We further identify robustness effects: opposite-signed errors can cancel, vanishing-tail link functions (e.g., logistic) stabilize clear-cut cases via saturation, and extra noise smooths belief maps and reduces their maximal local sensitivity. Simulations and real-world benchmarks show that EA-RAM is robust to Dual Error and achieves a better cost–performance Pareto frontier than centralized baselines, with additional gains from provider-side local information, validating its practical effectiveness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.