acceptodds
Under review as a conference paper at ICLR 2027

Forward Online Learning Deep Networks

Abstract

Modern deep neural networks are typically trained offline via computationally intensive large-scale backpropagation on massive datasets. In deployment scenarios, adapting the model weights to new data streams again requires costly backpropagation. In this paper, we derive from first principles a novel deep architecture, which we call FOLD, that supports both efficient *offline training via backward propagation* and *online learning via forward propagation*. Each layer consists of *a novel linear-time operator* that can replace both the attention and the MLP operators in a Transformer. The weights of each such operator are essentially a learned dictionary for its input data; in this way online adaptation to incoming new data becomes an online dictionary learning problem. As a result, the model weights can be updated at test time through *forward propagation as an optimization step* to better align the dictionaries with every new sample sequence. Experiments on synthetic and real-world benchmarks show that this new architecture performs competitively against the state-of-the-art linear-time recurrent deep architectures such as Gated DeltaNet with back propagation training while demonstrating qualitatively superior online learning capabilities with its forward propagation mechanism. We also show that incorporating explicit forward propagation into the architecture has side benefits, such as strong long-context performance with only subquadratic operations. Our results call into question the necessity of solely using back propagation to adapt pretrained models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.