Probe Before You Route: LLM Selection via Test-Time Probing
Abstract
Large language models (LLMs) vary substantially in capability and inference cost, making it important to select an appropriate model for each query. Conventional LLM routers select a model before any candidate model generation, potentially lowering routing performance because they overlook informative signals in the generated tokens. We therefore propose \bf Probe-then-Route, which first probes each candidate LLM by allowing it to generate a small number of output tokens for the query and then selects one model based on the query together with the generated tokens. To reduce the additional inference cost of probing every candidate, we further introduce a variant, dubbed \bf Selective Probe-then-Route, that first selects a candidate subset, probes only those candidates, and routes among them using the same trained router without retraining. We evaluate both methods using eight heterogeneous LLMs on 12 benchmarks spanning six domains, with one in-distribution (ID) and one out-of-distribution (OOD) benchmark per domain. Probing all eight candidates with up to eight output tokens per candidate improves average ID accuracy by 2.73% over the strongest baseline and outperforms all baselines on three of six ID benchmarks. \bf Probe-then-Route also generalizes effectively out of distribution, achieving the highest accuracy on two of six OOD benchmarks, including one tie, and remaining competitive on the others. With \bf Selective Probe-then-Route, additional inference cost can be reduced by 56.4% on ID benchmarks while retaining 83.3% of \bf Probe-then-Route's accuracy gain over the strongest baseline and still outperforming all baselines across three ID benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.