Functional Prox: Decoupling Importance Estimation from Optimization for Continual Learning in Spiking Neural Networks
Abstract
Spiking neural networks (SNNs) offer a promising path toward energy-efficient intelligent computing, yet retaining prior knowledge while continually learning new tasks remains a key challenge for continual adaptation. To mitigate catastrophic forgetting, importance-based regularization methods typically use gradients to estimate parameter importance and consolidate prior knowledge by constraining changes to important parameters. However, the non-differentiability of spike generation makes these estimates depend on the choice of surrogate gradient. Even when network parameters and forward predictions remain unchanged, changing only the backward surrogate can substantially alter hidden-parameter importance rankings or even cause importance values to explode, resulting in misplaced or excessive parameter constraints. To address this issue, we propose Functional Prox, a continual learning method that decouples functional importance estimation from surrogate-gradient optimization. Functional Prox estimates importance solely through forward spiking computations, measuring changes in predictive distributions under channel-level interventions to quantify functional contributions. These scores are converted into layer-normalized protection weights and combined with a closed-form, non-expansive proximal update to control constraint magnitudes and parameter displacement. Experiments on CIFAR-100, Tiny-ImageNet, and ImageNet-100 under class-incremental learning (CIL) setting show that Functional Prox significantly outperforms EWC and other representative regularization methods, improving both knowledge retention and new-task learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.