Decoupled Polynomial State-Space Iterations for Matrix Polar Approximation, with an Application to Muon
Abstract
We introduce Gram-matrix-based iterations (GMBI), a decoupled polynomial state-space parameterization that separately controls propagation of a square auxiliary state and accumulation of the rectangular output under a restricted matrix-computation budget. Building on the finite-depth polynomial framework of Polar Express, GMBI supports both compact short graphs and robust full-restart graphs. For restart graphs, we freeze a guarded two-layer prefix, materialize its outputs from FP16 calibration runs to rebuild true Gram states, and optimize the suffix by scalar rollouts from the realized restart spectra, adapting to prefix-induced distribution shifts and finite-precision effects. In controlled FP16 experiments, the decoupled GMBI parameterization outperforms a learned tied-polynomial control under both robust-only and Muon-aware fitting protocols, while the Muon-aware protocol provides an additional improvement for the decoupled variant. Tests on unseen spectra, matrix shapes, and held-out GPT-2 Medium updates show that these gains are not confined to the calibration set, while restart controls separately identify the benefit of finite-precision Gram-state repair. Kernel benchmarks show that compact GMBI graphs can reduce polar-call latency on large rectangular matrices under both compiled and specialized backends. Multi-seed GPT-2 Small training with Muon matches Direct PE in final validation loss while achieving modest reductions in mean step time, showing that the matrix-level gains can survive end-to-end training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.