Certifying LLM Serving Updates with Cost-Optimal Evaluation
Abstract
When an LLM serving pool changes, an operator must decide whether to retain an incumbent, replace it with a candidate, or combine them in a cascade. Evaluating all models on common queries can itself be costly. We formulate this pre-deployment decision as fixed-confidence identification with costly evaluation actions, separating the serving configuration to be selected from the observations used to select it. The characteristic certification cost quantifies the minimum cost of distinguishing the true environment from alternatives that change the optimal serving decision. It yields an instance-dependent lower bound, which our TRACK-AND-CERTIFY procedure asymptotically attains as under finite-class, stationary, identifiable observation models with known costs and a unique true optimum. For LLM serving, replacement depends on marginal performance, whereas a cascade triggered by verified incumbent failure also depends on same-query rescue behavior. Controlled replay across 11 model pools shows that richer multi-candidate feedback reduces certification cost only on selected pools, and that the characteristic-cost ratio strongly predicts this realized benefit. Decision-aware singleton allocation reduces median certification cost by relative to uniform singleton evaluation. These results motivate designing model evaluations around the serving decision they must resolve.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.