Scaling Zero-Order Pretraining Through Model Sharding
Abstract
Zero-order optimization (ZO) enables training without backpropagation, or retention of activations, making it relevant to forward-only hardware and non-differentiable loss, but its gradient error grows with perturbed dimension. This inhibits large model training. Sharded Optimization Mixture of Assemblies (SOMA) is an architecture designed with ZO in mind. SOMA is an ensemble of LSTM experts that train independently on clusters of data using simultaneous perturbation stochastic approximation (SPSA). Its separable loss function removes cross-expert perturbation noise at the cost of jointly learned representations across domains. Experts train independently, without exchanging gradients, activations or optimizer state. We use 80,000 estimated RTX 5090 GPU-hours to study SOMA compared to baseline methods. Modest sharding improves training compute efficiency over all tested monolithic ZO controls. We study a 8.44M model at 150 aggregate GPU-hour budget and show SOMA with 64 perturbations reaches 1.76 test nats/byte, versus 2.00–2.11 for monolithic SPSA at 64, 256 or 1,024 perturbations and 2.21 for EGGROLL. On WikiText-103, these frozen checkpoints reach 2.07, 2.25–2.36 and 2.49, respectively. We isolate the mechanism and prove that independent losses reduce relative gradient error to approximately of a shared-loss estimator's. Holding starting weights, data, perturbations and compute fixed, local rather than summed losses lower SOMA test loss by 0.035 nats/byte after 1,000 updates across three seeds. Finally, we show larger ensembles offer a separate inference benefit. At similar model size with top- routing (), SOMA achieves 2.36M tokens/s versus 257k for SOMA (, including routing), at lower test loss (1.68 versus 1.71), albeit with SOMA using as much aggregate training compute. We release all training and evaluation code and checkpoints for reproduction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.