Learning (and Forgetting) as Storage and Compression
Abstract
Dep neural networks exhibit strong generalization, yet the mechanisms underlying feature learning and its failure to learn under distributional shift (catastrophic forgetting) remain poorly understood. We introduce the Storage, Compression, and Targets (SCT) framework, which unifies neural learning as the interplay between storing associations, compressing representations, and aligning neural activity with local targets. We show that classical credit assignment algorithms—including backpropagation, predictive coding, and target propagation—can be formulated as specific instances of SCT rules. Through this lens, we identify compression as a primary driver of feature learning: once associations are stored, compression alone can reorganize their retrieval geometry, producing feature learning and generalization without further storage. However, the same mechanism can also drive forgetting. By studying the representation drift due to storage and compression rules, we prove that fully preventing task degradation requires corrections coupled to the current sample and, once representations reorganize, access to past inputs through replay or a generative proxy. Finally, we empirically validate our findings by deriving a replay-based correction mechanism that improves retention and sample-efficiency of replay-based methods for class-incremental learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.