Linear-to-Nonlinear Transitions in the Learning Dynamics of Recurrent Neural Networks
Abstract
How do recurrent neural networks (RNNs) acquire the temporal computations required by a task? We use the best causal linear (BCL) predictor—determined by input and target second-order statistics—as a task-defined reference for answering this question. First, we connect the BCL predictor to its minimal recurrent realization, providing an exact functional benchmark and task-defined poles. Second, we show that an overparameterized linear RNN realizes this predictor through a small set of residue-carrying modes while most of its spectrum remains behaviorally weak. Recurrent gain accelerates learning only up to a near-marginal regime, where predictive modes approach the BCL poles as nearby unused modes move away. Third, on a nonlinear memory task, a nonlinear RNN reaches BCL-level performance before improving beyond the optimal-linear loss and forming stable memory states around the same slow scaffold. A controlled filtered-cubic task shows that this linear-first ordering depends on the training regime. The BCL thus links task statistics, recurrent dynamics, and the emergence of nonlinear computation during learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.