acceptodds
Under review as a conference paper at ICLR 2027

The Geometry of Forgetting in nGPT: Disentangling Weight Rotations and Scale Dynamics

Abstract

nGPT achieves high training efficiency by constraining its major vector-valued parameters to unit hyperspheres while learning explicit scale parameters. This parameterization naturally decomposes parameter updates into weight rotations and scale changes. However, when a pretrained model is further trained, it remains unclear which of these geometric changes harm the behavior learned during pretraining. We study this question in domain-adaptive continued pretraining, analyzing which kinds of rotations and scale changes drive forgetting. We show that, at matched rotation angles, rotational damage varies strongly with alignment to task-sensitive directions. At matched displacement sizes, scale damage varies with the spread of per-token logit reweighting. Based on these channel-specific mechanisms, we design a two-term regularizer and use it to intervene on the two channels separately. Weight rotations drive most net forgetting, while scale damage is largely offset by output-projection interactions. With both penalty terms, suppressing task-sensitive rotations exposes this scale damage, which the scale penalty then controls. nGPT's hyperspherical parameterization thus provides not only a mechanism for training efficiency but also a coordinate system for understanding and controlling change during continued pretraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.