Maglev: Improving the Recommendation Scaling Laws via Unified Scaling
Abstract
Industrial recommenders face diminishing returns from dense scaling, while sparse and depth scaling make global feature interaction and distributed training pro- hibitively expensive. We introduce Maglev, a model–system co-design that decom- poses a recommender into hardware-local pipeline stages. It combines (1) Maglev Indexers, which learn capacity-bounded stage inputs, as soft projections within one scale-up domain or hard feature partitions across them; (2) Maglev Induct, which carries a low-rank residual across stages; and (3) Maglev Rail, a pipeline schedule that stays bubble-free with synchronous gradient clipping. We conduct an industrial-scale study that jointly scales sparse capacity and dense compute (28B samples, 700 features, four tasks). Applied to Wukong and DHEN, the strongest of 10 screened backbones, Maglev improves the scaling law along every axis: every Maglev configuration outperforms the best baseline configuration at equal or lower compute and sparse capacity, and the lead widens as sparse capacity and compute grow. The full system trains up to 1.3× faster than hybrid parallelism. On three public benchmarks, Maglev improves five further backbones in all 15 matched- compute settings. Maglev has been launched in three flagship recommenders at a major social media company, improving RLL by up to 0.4% in A/B tests.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.