The art and principle of universal scaling
Abstract
Scaling laws have become the bedrock for foundation model training, yet standard scaling laws are largely restricted to evaluating cross-entropy loss at terminal checkpoints – leaving the continuous dynamics of general observables under-explored. In this work, we establish universal scaling laws across large language models (LLM) and multi-modal models (MMLM) that accurately characterize the any-time-any-scale trajectories of losses, evaluation metrics, and parameter norms. To establish this unified framework, we integrate three key pillars: conventional end-point scaling laws, convex optimization theory, and an enhanced collapse transformation. We systematically validate this framework across diverse model architectures, optimizers, training hyperparameters, compute regimes, and modalities, providing a principled recipe for universal scaling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.