Controlling Symmetry Drift in Forward-Only Continual Learning
Abstract
Edge devices need to adapt in the field despite limited access to backpropagation, making node perturbation an appealing alternative based on local activity, neuron-level perturbations, and scalar loss feedback. Yet persistent perturbation noise can drive parameter growth over long data streams and undermine a learner's ability to adapt to later tasks. We show that, to leading order in a frozen-state analysis, non-gradient perturbation noise produces a mean-zero diffusion along loss-invariant rescaling symmetries, whereas gradient updates have no first-order component along these directions. We introduce gauge-damped node perturbation (), which periodically corrects symmetry-charge displacement through local reciprocal rescaling of each hidden neuron's incoming and outgoing weights. For positively homogeneous activations, each positive rescaling preserves the current function exactly, making the correction prediction-unbiased; the unclipped update contracts the targeted charge for an isolated layer pair, while the guarantee depends on the architecture. Across continual-learning streams, substantially controls charge displacement on long ReLU runs and can improve late-task accuracy; tanh experiments and simulated quantized writes further show that the benefit depends on the activation and write protocol.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.