Rank-1 Fisher for Continual Learning: When Is Its Geometry Faithful
Abstract
Mean-gradient rank-1 empirical-Fisher surrogates offer a memory-efficient approach to continual learning, but their practical utility does not establish that one retained direction faithfully represents loss-gradient second moments on unseen data. We examine this geometric interpretation theoretically and empirically. The motivating linear-model proof omits an input outer product: the gradient is proportional to , rather than . Under the corresponding linear–Gaussian model, we derive the exact population second-moment operator and show that, although vanishes with dimension for a scaled-identity residual, the optimal rank-1 relative Frobenius error approaches . Simulations further show that same-sample evaluation can spuriously favor rank-1 at low , whereas this advantage disappears on disjoint test samples. Held-out evaluations on three frozen masked diffusion language models spanning 219M to 8B parameters reveal no uniform rank-1 advantage: at the largest calibration size, rank-1 slightly outperforms diagonal on SMDM-219M, but underperforms on SMDM-1.14B and in 14 of 15 LLaDA-8B settings. An exact error decomposition gives a diagnostic answer: held-out faithfulness jointly requires rank-1 capacity, transfer of the mean-gradient direction, and transfer of its fitted scale. A source-setting check on small MNIST UNets provides a favorable local boundary, but the overall results vary across objectives, scales, layers, and corruption levels. Mean-gradient rank-1 Fisher therefore need not provide a faithful held-out reconstruction, although this does not rule out its consolidation utility. Code is available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.