To Gate or Not to Gate? When Does an Input-Dependent Gating Pay Off in Multi-Class Dendritic Gated Networks
Abstract
Dendritic Gated Networks (DGNs) consist of neurons that each have several weight vectors (their branches), each paired with a fixed random hyperplane. The input, together with the hyperplanes, determines which branches are active, contribute to the output and get updated. This mechanism makes DGNs resistant to catastrophic forgetting, able to approximate arbitrary nonlinear functions while being convex in their weights, and trainable without backpropagation. DGNs are binary classifiers, and existing multi-class extensions support classes one of three ways: training one full network per class (one-versus-all), training networks over an efficient binary code, or applying a shared gated expansion, in which one fixed set of random half-space gates is applied once per input, the per-branch blocks are concatenated into a single feature vector, and independent Bernoulli readouts read that vector. We propose training those readouts jointly by softmax regression: the learning objective stays jointly convex, the update stays local up to one shared normalizer per sample, and the design keeps the expansion's one to two orders of magnitude fewer parameters than the one-versus-all formulation, and its single expansion in place of the forward passes, which is what makes tractable. We then investigate when using input-dependent gates pays off, against controls matched in dimension, parameter count, update rule and training protocol. Under domain-incremental training with smooth image rotation ( and ) and pixel noise (), our proposed method outperforms the dense, static-gate and linear controls on all-task and first-task accuracy. Given the same frozen features, it also outperforms the training-free nearest-class-mean classifier, and against and over all tasks under rotation and noise. With a replay buffer on split-CIFAR-100 it also outperforms a backpropagation-trained network with Dark Experience Replay++ (DER++) at every buffer size we ran, against at rows. We attribute the gating margin to regime structure, where the input-label mapping differs between regimes that can be identified directly from the input itself. In a class-incremental problem, which has no such regime structure, our method's margin over its static-gate control shrinks to the margin it already has under i.i.d. training, as predicted. It still outperforms one-versus-all and, storing no past data, reaches against for a backpropagation-trained network with Learning without Forgetting, but an untrained dense random projection of the same width matches it and the nearest-class-mean classifier leads it ().
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.