Deployment-aware model selection
Abstract
Two serving stacks can agree on every number a leaderboard publishes and still be the right and the wrong choice for the same deployment. Of Open LLM models, are the best available choice at some task mixture while topping none of the seven numbers published. A system's risk profile, its best achievable risk at each operating condition, is the minimal object an evaluation can depend on; practice puts point masses on it, and we characterize the linear and convex rules completely. What decides when that choice matters is carried not by a profile but by the differences between candidates'. Aggregating over a range ranks a pair exactly as scoring it at the range's average does, precisely when that difference profile is affine. Shared curvature therefore cancels, and profiles may bend arbitrarily yet reorder nothing. No regular linear rule reading finitely many conditions is deployment-complete, and we localize what a suite of covering radius can miss. The boundary is also a noise-calibrated test of whether a full-profile criterion can change a selection. It is correct on all eight systems we try, and bending predicts divergence, not benefit: four bend and two pay, by – on re-steerable RLHF pipelines and on moderation. A replication across hospitals shows that advantage is largely a hedge against estimation error, fading as profiles improve and the rules converge ( agreement to ). What pays throughout is evaluating at the deployment's own condition, – away in total variation from the headline's, using scores already published.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.