Append-Only Mixture-of-Experts: Efficient Training and Flexible Deployment
Abstract
We introduce Append-only Mixture of Experts (MoE), which grows an MoE during training and keeps each smaller model available for deployment. We train a small base model, then add expert rings while freezing existing parameters. Each ring has its own router, so disabling later rings recovers an earlier model. We train four sizes, from 199M to 1.597B parameters, using solution traces for String Algorithms, Maze Navigation, Arithmetic Sums, and Conway's Game of Life. At 35,000 updates, the four-stage run uses 6.80 hours of recorded training duration on four H200 GPUs, compared with 21.18 hours for the tested full-size scratch baseline. Mean final-answer accuracy is 99.22% and 80.27%, respectively. All six checks of earlier prefixes produce bitwise-equal logits. Having this adjustment in size also affects decoding speed, where 199M base decodes at 209.4 tokens/s, compared with 39.8 tokens/s for the largest.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.