acceptodds
Under review as a conference paper at ICLR 2027

MarCap: From Marginal Routing Probabilities to Risk-Bounded MoE Workspace Provisioning with Lossless Replay

Abstract

MoE kernels reserve workspace before routing loads are known. Reserving for the worst case wastes memory, while fixed capacity factors fail to quantify overflow risk. We present MarCap (ginal-based acity provisioning), which maps routing marginals and a user-specified overflow risk budget to per-rank receive-buffer capacities. It optimizes over all possible fixed top- joint routing laws consistent with the marginals, covering arbitrary within-token expert co-selection correlations. Under the IID token assumption, we formulate receive-buffer capacity bounds in terms of extremal moment-generating functions and use exact feasibility conditions to reduce their computation to a small linear program that is solved efficiently by a lower convex-hull construction. We validate MarCap on 2.737B token-level routing records from two LLMs. Analysis identifies inter-token dependence in prefill and marginal drift as sources of potential undersizing. Traffic mixing reduces dependence-related shortfalls, while sliding-window recalibration mitigates drift-related undersizing. These results clarify the theory's practical scope. We implement MarCap in \mathtt{megamoe\underline{\}replay}, our fused MoE operator with bounded workspace and lossless replay. Across 660k EP4 invocations on Ascend 950PR, \mathtt{megamoe\underline{\}replay} preserves reference outputs and reduces scratch workspace requirements by 44.1–71.4% versus full reservation across the evaluated configurations. In the tested decode workloads tested on hardware, which are well aligned with the IID model, MarCap keeps the observed frequency of overflow which would trigger additional replay rounds within the user-specified risk budget. Experimental data confirm the resulting benefit: substantial scratch workspace savings with mean latency comparable to the worst-case reservation, with a median absolute difference of 0.36% across the tested decode configurations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.