acceptodds
Under review as a conference paper at ICLR 2027

Perturbation Propagation: Scalable Local Learning without Weight Transport

Abstract

The growing energy demands of artificial neural networks on digital hardware increasingly limit the level of intelligence they can reach, motivating the search for potentially far more energy-efficient analog implementations. The brain offers an existence proof that a physical system can learn complex tasks at very low energy cost, but a biologically plausible learning algorithm that trains at scale remains elusive. Backpropagation is not biologically plausible: it relies on exact derivatives from mathematical models of neuron responses and on weight transport, which copies forward weights into the backward pathway. More biologically plausible alternatives struggle to match backpropagation's scalability and algorithmic simplicity. We propose Perturbation Propagation (PerProp), which keeps the forward-backward structure of backpropagation but replaces both operations with local measurements: neuron derivatives with slopes measured by small forward perturbations, and weight transposes with feedback matrices learned from random probes. Because each neuron acts on its own input, one pair of perturbed evaluations measures a whole layer's slopes; because weights are shared across examples, feedback matrices are learned once per update. In Transformer pretraining, PerProp scales at a rate comparable to backpropagation with model size and compute, and at a faster rate with data, batch size, and depth. On GPT and Mamba-2 language models, it outperforms the feedback-alignment and perturbation-based alternatives we evaluate, both in and out of distribution. On spiking, comparator, and probabilistic-bit networks, it rivals or exceeds surrogate-gradient training without hand-designed surrogate derivatives. On analog hardware, PerProp trains physical neural networks directly, learning entirely on the BrainScaleS-2 neuromorphic chip and substantially outperforming reward-modulated STDP. Component-level estimates for analog crossbars further project orders-of-magnitude lower training energy than GPU backpropagation, an advantage that grows with network size. PerProp thus offers a unified approach to scalable, local learning in digital and physical neural networks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.