acceptodds
Under review as a conference paper at ICLR 2027

CCQ: Compensation-Conditioned Quadratic Selection for Absorbing Skipped Experts in MoE LLMs

Abstract

Executing only half of the routed experts per token halves the routed-expert work of a fine-grained Mixture-of-Experts (MoE) model, but the absent experts leave an error. We split it into a common mode, the mean expert output scaled by the absent routing mass, and an expert-specific remainder. Restoring each part exactly on three MoEs shows that both matter, in model-dependent proportions, and that their costs nearly add, as a second-order analysis of the KL divergence predicts when their interaction is small. We therefore compensate them with different tools. A mass-gated low-rank readout of the always-executed shared expert, trained by KL distillation, recovers what the shared expert predicts of the absent output. Diagnostics show it is not merely a common-mode estimate, yet on the model where the common mode dominates, it removes as much KL as the exact common mode. The readout leaves most of the specific part, but each routed expert's output energy is predictable. We skip the half with the least predicted cost and add a sketch of the costliest absent expert at an eleventh of its width, trained on the residual the readout leaves. Implemented in vLLM with CUDA kernels, on Moonlight-16B-A3B, DeepSeek-V2-Lite and Qwen3.6-35B-A3B, the method, CCQ, lowers the full-vocabulary KL to the full model by 23–80% against skipping alone and by 13–18% against a LoRA with the same distillation budget, the strongest of seven baselines. Selection and sketch provide this gain; the readout alone only matches the LoRA. The method retains 98.0–99.2% of the full model's average score on 12 benchmarks, though mostly not significantly different from the readout or the LoRA. In vLLM's throughput benchmark on one GPU, it generates 1.19–1.23× as many tokens per second as the full model at batch size 1, up to 1.48× at batch size 4 and 1.06–1.15× at 256, mostly owing to skipping itself.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.