Layer-Selective Utility Gating for Continual Learning
Abstract
Neural networks in continual learning suffer from plasticity loss. Utility-based perturbed gradient descent (UPGD-W) mitigates this problem through utility-dependent modulation of parameter updates. We investigate how this modulation should be allocated across network layers. We first analyze an idealized continuous-time model and show that, under regularity and stationarity assumptions, its expected displacement-based utility rate relates to gated local curvature with an explicit correction for non-equilibrium probability currents. We then derive a distributional tracking bound that separates the contributions of representation and head drift. These results provide a theoretical motivation for comparing output-only, hidden-only, and full gating across different adaptation settings. On three supervised label-permutation benchmarks, output-only gating applies utility modulation to only 0.16%–2.5% of parameters while achieving 114%–285% higher mean accuracy than Shrink and Perturb and up to 13% higher than full gating. All layers continue to update. The gating variants perform similarly on Input-Permuted MNIST. In Gridworld with controlled dynamics shifts, hidden-only gating achieves higher mean return than output-only gating. Together, these findings show that the benefits of utility gating depend on its placement and the adaptation demands of the learning setting.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.