acceptodds
Under review as a conference paper at ICLR 2027

When Routing Edits Do Not Add Up: Budgeted Selection in Mixture-of-Experts Models

Abstract

Compressing a sparse mixture-of-experts model changes which experts its router selects. Candidate edits are often scored one at a time, although an upstream edit changes the hidden states on which later edits act. We find that edits that help in isolation can lose reference-answer log probability when combined. Under additive logit responses, a local quadratic expansion expresses pairwise interaction through the softmax Fisher inner product. We propose Ranked Validation, which measures the empty set and every single edit, orders larger candidate sets by summed single-edit scores, measures the leading sets, and returns the best set measured. For bounded interactions, we establish a sharp worst-case regret limit for deterministic scan-only selection; Ranked Validation satisfies the same upper bound once its record contains the additive maximizer. Across four checkpoints and thirteen candidate populations, it improves pooled exact selection over greedy and random validation under matched query budgets. On the half-pruning population, four extra evaluations recover 29.0% of the lost reference-answer log probability. A set selected once and reused across held-out questions adds little beyond a calibration-selected single-edit baseline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.