Byte Economics: Million-Token Inference of a 2.8-Trillion-Parameter MoE on a Consumer Desktop
Abstract
Mixture-of-experts language models now exceed a trillion parameters, while the largest consumer GPU holds 32 GB. We show that serving such a model on a desktop is governed by byte economics: decode time is the bytes each token moves across the memory hierarchy divided by the bandwidth of the links, and prefill time is set by how many times the expert store is read. We build OMOE, a tiered serving system that runs the native 2.8-trillion-parameter Kimi-K3 checkpoint (1.4 TB of MXFP4 experts, unmodified) on two RTX 5090s, 192 GB of DRAM and a 3.2 TB storage tier, with outputs gated byte-exact or within of a full-VRAM reference. A bandwidth cost model with no fitted constants guided four system generations from 4.62 to 1.65 s/token at short context and places the software ceiling for this model on this machine at 0.58 s/token. A layer-major prefill schedule reads each expert at most once per prefill, so prefill throughput falls only from 69 to 47 tokens/s between 96K and one million tokens of context, where needle retrieval passes with 6.1 h to first token and 3.53 s/token decode. The instrumented deployment measures what a production MoE asks of a memory hierarchy: an LRU hit rate of 0.34 at 3% of the store, 3.0 TB written per million-token prefill, and the bytes a model designer could remove. Concurrent consumer-scale systems report contexts up to 65K; the system, telemetry and acceptance harnesses will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.