acceptodds
Under review as a conference paper at ICLR 2027

A Thresholded ReLU Unit with Fixed Output Weight Goes Dormant under Target Drift: Phase Diagram and Entry Time

Abstract

Networks tracking moving targets lose ReLU units to dormancy: units active on almost no input get almost no gradient. Lacking a target-speed parameter, existing dormancy theory explains why such units stay dormant, not when they become dormant. We study one ReLU unit with a learned threshold and fixed output weight, trained online on a teacher whose weight vector random-walks on the sphere at fixed speed. In the large-dimension limit, its weight length, teacher overlap and threshold obey three differential equations with closed-form resting states. We map the boundary between tracking and dormancy over drift speed and learning rate: at slow drift the tracking state collides with an unstable state and both vanish; at fast drift it survives but is dormant. With nothing fitted, the equations predict where and when dormancy sets in, and simulations match once finite-dimension fluctuations are kept. These equations give the mechanism: drift keeps an error alive, at rates causing dormancy each update's squared length inflates the weight vector, so the threshold must outgrow the weight length. Since this inflation vanishes with the learning rate, we prove an isolated unit in the large-dimension equations needs a rate bounded away from zero for dormancy. Yet one of two units sharing one output error goes dormant where the isolated unit never does. In both settings a trained output weight settles at a small scale and prevents dormancy everywhere tested.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.