acceptodds
Under review as a conference paper at ICLR 2027

Few-Shot Model Onboarding for LLM Routing

Abstract

Model routing reduces the inference cost of large language models (LLMs) by sending each prompt to the cheapest adequate model in a pool. Frequent model releases and costly response labeling motivate onboarding from a handful of labels, yet this setting remains largely unexplored. Recent routers can admit unseen models without retraining, but they either calibrate on a large set of randomly drawn labeled prompts or train an additional item response model to select prompts. We identify prompt selection and label extrapolation as two key factors in few-shot calibration. We propose Active Anchor Calibration (AAC), which clusters existing models' responses to select a few anchor prompts, and uses a closed-form ridge map, together with a cost prior, to extrapolate their labels to model quality and per-prompt predictions. Under a factor model with known prompt factors, we bound pool-level routing regret through an anchor design criterion. Experiments on four public benchmarks show that, with 5 labeled prompts per new model, AAC recovers 88–93% of full-calibration routing quality, against 74–83% for random calibration, which needs about 40 labels to match it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.