RankShift: Dual-Reference Identification of Routing Value under Delayed Feedback
Abstract
Aggregate routing gains do not identify routing value. A positive gain may come from matching experts to states or simply from calling a globally stronger expert more often. It may also hide local losses against the deployed model. We therefore define operational routing value relative to two fixed references. The deployed anchor tests current-system improvement, and a same-budget state-independent policy tests value beyond a registered blind allocation. A single comparison need not determine both properties. RANKSHIFT authorizes a fixed proposal only when matured paired outcomes give positive lower bounds against both. Under explicit dependence and drift conditions, we describe the delay and support required for this decision. On a held-out language-model routing benchmark, one validation-frozen rule authorizes three of twelve proposal learners at the primary seed without retraining. Their gains are +0.0224, +0.0227, and +0.0237 with no observed local fallback regression. The other nine remain on the anchor. An independent routing trace plus delayed replay and three non-LLM domains test the same mechanism.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.