acceptodds
Under review as a conference paper at ICLR 2027

Where Does Harmful Replacement Come From? Coverage and Switching in Frozen Model Routing

Abstract

Adding a frozen expert gives a router another possible answer, but also another way to replace a correct incumbent answer with an incorrect one. We introduce a diagnostic protocol for this trade-off when two new candidates are available. Holding a fixed support budget of 64 candidate evaluations per intervention, a pair-fixed policy uses one candidate on every eligible request, whereas a per-request policy chooses between the two candidates separately for each request. Across two development assets, per-request choice increases net accuracy gain relative to the incumbent by 0.182 percentage points (pp) compared with pair-fixed choice, while increasing harmful-replacement mass by 0.201 pp. An exact request-wise decomposition attributes the difference to coverage, where one policy returns a candidate and the other retains the incumbent, and switching, where both return candidates but return different ones. A controlled experiment shows that these decisions interact: expanding candidate use can help or hurt depending on which candidate is returned on the added requests. A held-out text-model asset and an independently constructed score reproduce the interaction and the sign reversal of the coverage effect, whereas switching-harm reduction is not stable across assets or score constructions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.