acceptodds
Under review as a conference paper at ICLR 2027

Drift Is Not Enough: A Diffusive Floor for Gradient Propagation in Recurrent Networks

Abstract

Whether a recurrent model learns long-range dependencies is determined by whether its gradients survive backpropagation over long horizons without exploding or vanishing. This behavior is governed by the top Lyapunov exponent , which characterizes the asymptotic rate of gradient growth or decay. While recent regularization methods seek to ensure marginal stability by driving this exponent to zero, the dynamics governing the non-autonomous gradient propagation remain insufficiently characterized. In this work, we model gradient dynamics as a linear cocycle over an ergodic Markov chain. Within this framework, we prove that the Lyapunov exponent spectrum of Recurrent Neural Network (RNN) gradient dynamics is well-defined and independent of the input realization, acting as a deterministic Law of Large Numbers drift. Crucially, while this drift rate governs the asymptotic scaling, gradient transport over any finite horizon exhibits second-order fluctuations governed by a companion Central Limit Theorem with asymptotic variance . Consequently, eliminating drift () does not prevent degradation, but leaves a superpolynomial, two-sided dispersion scaling as , rendering vanishing and explosion equiprobable. We analytically prove that the asymptotic variance is strictly non-degenerate () for any choice of recurrent weights in vanilla RNNs with saturating activations, showing that gradient dispersion is an inescapable feature of these models. In our experiments, penalizing both the drift (gradient flossing) and the variance reduces empirical diffusion by 63–74% and improves psMNIST accuracy from (unregularized) and (flossing alone) to . Because this formulation requires only a recurrent state, it extends beyond vanilla RNNs, and we observe the same growth in gated, orthogonal and selective state-space recurrent models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.