acceptodds
Under review as a conference paper at ICLR 2027

Recalibrated Routing for LLM Judges: Optimizing Classification Metrics On A Budget

Abstract

Routing is important in LLM-as-a-judge research because it combines models with different comparative strengths in a cost-effective manner. In practice, performance of such routers is often measured using interpretable classification metrics such as precision, recall, and score; however, existing methods typically optimize objectives in the form of average proxy quality that need not align with these deployment metrics. In this paper, we develop a budget-aware framework that directly optimizes the chosen metric. In particular, we use a quasi-linear representation to formulate routing as a mixed-integer linear program with nuisance parameters. Furthermore, we apply a post-hoc calibration that guarantees to improve on the raw AI outputs with high probability. For theoretical guarantees, we characterize the population-optimal router and recalibration cutoff and derive regret bounds for our metrics in terms of nuisance regression estimation error. We apply the method on a real-world safety violation detection problem at a tech company and show that our routing method outperforms existing methods in the literature.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.