acceptodds
Under review as a conference paper at ICLR 2027

Scaling Zero-Order Pretraining Through Model Sharding

Abstract

Zero-order optimization (ZO) enables training without backpropagation, or retention of activations, making it relevant to forward-only hardware and non-differentiable loss, but its gradient error grows with perturbed dimension. This inhibits large model training. Sharded Optimization Mixture of Assemblies (SOMA) is an architecture designed with ZO in mind. SOMA is an ensemble of LSTM experts that train independently on clusters of data using simultaneous perturbation stochastic approximation (SPSA). Its separable loss function removes cross-expert perturbation noise at the cost of jointly learned representations across domains. Experts train independently, without exchanging gradients, activations or optimizer state. We use 80,000 estimated RTX 5090 GPU-hours to study SOMA compared to baseline methods. Modest sharding improves training compute efficiency over all tested monolithic ZO controls. We study a 8.44M model at 150 aggregate GPU-hour budget and show SOMA with 64 perturbations reaches 1.76 test nats/byte, versus 2.00–2.11 for monolithic SPSA at 64, 256 or 1,024 perturbations and 2.21 for EGGROLL. On WikiText-103, these frozen checkpoints reach 2.07, 2.25–2.36 and 2.49, respectively. We isolate the mechanism and prove that independent losses reduce relative gradient error to approximately of a shared-loss estimator's. Holding starting weights, data, perturbations and compute fixed, local rather than summed losses lower SOMA test loss by 0.035 nats/byte after 1,000 updates across three seeds. Finally, we show larger ensembles offer a separate inference benefit. At similar model size with top- routing (), SOMA achieves 2.36M tokens/s versus 257k for SOMA (, including routing), at lower test loss (1.68 versus 1.71), albeit with SOMA using as much aggregate training compute. We release all training and evaluation code and checkpoints for reproduction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.