CompareRoute: Learning Relative Expert Capability for Prefill-Based LLM Routing
Abstract
LLM routing selects an expert before answer generation, requiring reliable estimates of relative expert capability on each query. We propose CompareRoute, a prefill-based router that uses expert-conditioned dynamic layer aggregation to read multilayer states from a frozen Observer. A shared antisymmetric comparator predicts pairwise success-rate differences from the resulting expert-specific query representations. Training combines success supervision and continuous regret with a comparison loss that gives additional weight to pairs involving the router’s current top choice, proportional to their empirical success-rate gaps. Inference requires one Observer prefill and invokes only the selected expert for answer generation. Across three benchmarks with five heterogeneous experts, in-domain evaluation shows an average success-rate improvement of 1.06 percentage points over the strongest evaluated baseline. Selected-expert token usage decreases by an average of 30.65% relative to single-expert references, excluding Observer overhead. These results highlight the potential of CompareRoute to improve response quality while reducing token usage in multi-LLM deployments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.