How Much New Geometry Does Another Training Run Add? Effective-Dimension Scaling Across Aligned Parameter Trajectories
Abstract
Neural network training evolves in a high-dimensional parameter space, yet independent runs often exhibit shared geometric structure. Prior work has identified such structure in both converged solutions and training trajectories. How the joint spectrum of complete parameter trajectories changes with the number of independent runs remains less well understood. We organize checkpoints from multiple runs into a joint trajectory tensor and measure the effective dimension of its singular-value spectrum after symmetry-aware alignment and temporal centering. We compare the observed spectra with the weak and strong matched-rotation nulls and repeat the analysis across multiple nested trajectory orderings. Across the evaluated architectures and tasks, effective dimension grows rapidly over the first few runs, followed by diminishing marginal growth over the measured range of up to 40 runs. The observed dimension remains far below the strong matched-rotation null and the algebraic rank bound. These results show that aligned parameter trajectories have a strongly concentrated, effectively low-rank joint spectrum over the observed run-count range.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.