Sign-based optimizers are massless relativistic learners
Abstract
Sign-based optimizers are like photons: they travel at a constant speed and only change direction. This suggests a connection with special relativity. We derive Signum and Lion from relativistic learning dynamics with a *maximum speed of adaptation*, , for each parameter. The sign operation emerges as the massless (photonic) limit, while Lion's gradient mixing emerges from *curvature-aware dissipation* that damps motion more strongly along directions of high positive curvature. We further interpret equal- Adam as Signum with an adaptive mass determined by gradient variance. Combining this Adam-like adaptation with our finite-mass Lion derivation yields *Massive Lion*, which incorporates Lion's curvature-aware dissipation, variance-based mass adaptation, and a tunable finite rest mass. Its updates interpolate between an approximately linear, SGD-like regime and a saturated, sign-like regime. Experiments reveal the complementary roles of these mechanisms. On CIFAR-10 with ResNet-18, finite rest mass improves both sign-based and adaptive variants, bridging their gap with momentum SGD. Curvature-aware dissipation accelerates optimization on heterogeneous quadratic landscapes and improves language modeling. For a 12-layer transformer trained on 3.25 billion SlimPajama tokens, it consistently improves validation perplexity at high momentum across three learning rates. Our results show how physical reasoning explains relationships among existing optimizers, distinguishes the roles of their constituent mechanisms, and guides the design of new ones.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.