Decision Centric Active Learning for LLM Routing
Abstract
LLM routing selects, for each query, the model from an available pool that best trades off answer quality and cost. Predictive routers estimate these criteria for every model and pick the best one, which presumes training data in which queries have been evaluated on all models—data that is costly to collect and must be collected again whenever a new model is added. We study this model-onboarding setting: a router over a set of incumbent models is already trained, and a new LLM must be integrated using as few labeled queries as possible. Classical active learning reduces labeling cost by querying where a predictor is most uncertain, but it ignores the routing decision and therefore spends much of its budget on queries the new model would never win. We propose Adaptive Decision-Centric Query-by-Committee (ADC-QBC), which mixes, within every batch, standard query-by-committee samples that reduce global uncertainty about the new model with decision-centric samples on which the committee disagrees about whether the new model should replace the best incumbent, shifting the budget from the former to the latter through a decaying schedule. On the SPROUT benchmark, onboarding GPT-4o mini and Claude 3.5 Sonnet, ADC-QBC attains higher routing accuracy than standard query-by-committee in 54 of 72 combinations of label budget and cost sensitivity; when onboarding GPT-4o mini, with 4,000 labels it outperforms standard query-by-committee with 8,000 at intermediate costs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.