acceptodds
Under review as a conference paper at ICLR 2027

Neuromodulated Equilibrium Learning: Reward-Driven Learning from a Single Free Equilibrium

Abstract

Equilibrium Propagation (EP) is a biologically motivated alternative to backpropagation but requires free and nudged equilibrium phases, raising questions about the biological plausibility of phase-separated learning. We introduce Neuromodulated Equilibrium Learning (NEL), a single-equilibrium reward-based framework inspired by Attention-Gated Brain Propagation and neuromodulatory signaling. After reaching a free equilibrium, NEL propagates a reward-dependent modulatory signal derived from the selected action through reciprocal feedback connections to drive local synaptic updates, eliminating both the reward-nudged equilibrium and contrastive update. We evaluate NEL on MNIST, Fashion-MNIST, and CIFAR-10 using multilayer perceptrons and convolutional networks, with symmetric three-phase REP and standard EP as multi-phase baselines. NEL approaches three-phase REP in shallow architectures, with gaps of 0.79 and 1.31 percentage points on MNIST and one-hidden-layer Fashion-MNIST, while larger gaps of 3.13 and 6.21 percentage points emerge in the deeper Fashion-MNIST and CIFAR-10 settings. A direct cost feedback control using the full supervised output error also remains below three-phase EP, indicating that sparse reward information alone does not explain the performance gap between single-equilibrium and multi-phase learning. These results show that effective reward-driven learning can arise from a single free equilibrium while suggesting that computations introduced by additional nudged relaxations become increasingly important in more challenging architectures.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.