Beyond Orthogonality: Rotation Parameterization Shapes Sustained Learning
Abstract
Continual learning requires networks to keep learning as tasks change, yet prolonged training can erode this ability. Spectral constraints such as orthogonality can support adaptation, but leave open how the constrained weights should be represented and optimized. Motivated by rotation constructions shared by classical and quantum learning models, we examine this question while maintaining hidden-weight orthogonality throughout training. Across 1,000-task permutation streams on MNIST, CIFAR-10, CIFAR-100, and Tiny ImageNet, increasing rotation repetition from one to nineteen improves mean continual accuracy by 3.58–14.42 percentage points and increases the benefit of training history relative to networks trained anew on each task. Most of this improvement is reached by eight repetitions and persists under learning-rate selection and matched initialization; further repetition yields diminishing returns, while comparisons at a fixed parameter count reveal a complementary benefit from broader coordinate interactions. Beyond the number and arrangement of rotations, constraining their optimization coordinates also changes adaptation within the same nominal representable family. More broadly, these findings show that favorable weight spectra alone do not determine sustained learning: the structure and optimization of the underlying parameterization remain consequential.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.