The Memorization Cost of Generalization: A Quantitative Causal Law in Noisy-Trained Deep Networks
Abstract
Deep networks trained on noisy labels do two things at once: they learn generalization from the clean samples and they account for the noise samples by memorizing them. We show that these two processes occupy complementary channels of the representation's topology. Using a two-level persistence pipeline (Topo) that separates persistent homology into a within-class component (loops internal to a single class, the generalization channel) and a cross-class component (loops that close only across classes, the memorization channel), we establish that memory and generalization are causally separable: memory can be added, stripped, and — the central result — its cost on generalization can be billed quantitatively. Along the overlay path (freeze a clean-trained model, then unfreeze flipped samples with clean samples still present), the memorization cost obeys the law i.e., each 1% of memorized flipped samples costs validation accuracy, with the slope proportional to the noise:clean sample ratio . At the shared 11M reference capacity, the slope constant clusters near across datasets and class counts: (CIFAR-10), (SVHN), (CIFAR-100, 100 classes), and (VGG on SVHN); the per-setting spread is the image term of the displacement mechanism, not a capacity signature. The law also holds on real human label noise: on CIFAR-10N the calibrated constant predicts the measured memorization slope to a mean ratio of (3 seeds), while maximally class-confusable flips bound the coefficient's noise-structure sensitivity at . Across capacities varies systematically (0.470.34 over a width sweep) and is traced to a normalized clean-sample feature displacement, (Zhang et al., 2026, arXiv:2608.30487). Every calibration cell is 3-seed. The constant is a through-origin fit; a low-noise extrapolation to stays within on both datasets but with larger residual scatter, delimiting the law's cleanest domain to the mid-noise range . Memory, we have shown, is an invertible, stackable layer, not an uninterpretable black-box capacity: its cost to generalization obeys a quantitative causal law with a calibrated constant that clusters near across datasets and class counts at the shared reference capacity, with its capacity dependence mechanistically traced.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.