acceptodds
Under review as a conference paper at ICLR 2027

The Memorization Cost of Generalization: A Quantitative Causal Law in Noisy-Trained Deep Networks

Abstract

Deep networks trained on noisy labels do two things at once: they learn generalization from the clean samples and they account for the noise samples by memorizing them. We show that these two processes occupy complementary channels of the representation's topology. Using a two-level persistence pipeline (Topo) that separates persistent homology into a within-class component (loops internal to a single class, the generalization channel) and a cross-class component (loops that close only across classes, the memorization channel), we establish that memory and generalization are causally separable: memory can be added, stripped, and — the central result — its cost on generalization can be billed quantitatively. Along the overlay path (freeze a clean-trained model, then unfreeze flipped samples with clean samples still present), the memorization cost obeys the law i.e., each 1% of memorized flipped samples costs validation accuracy, with the slope proportional to the noise:clean sample ratio . At the shared 11M reference capacity, the slope constant clusters near across datasets and class counts: (CIFAR-10), (SVHN), (CIFAR-100, 100 classes), and (VGG on SVHN); the per-setting spread is the image term of the displacement mechanism, not a capacity signature. The law also holds on real human label noise: on CIFAR-10N the calibrated constant predicts the measured memorization slope to a mean ratio of (3 seeds), while maximally class-confusable flips bound the coefficient's noise-structure sensitivity at . Across capacities varies systematically (0.470.34 over a width sweep) and is traced to a normalized clean-sample feature displacement, (Zhang et al., 2026, arXiv:2608.30487). Every calibration cell is 3-seed. The constant is a through-origin fit; a low-noise extrapolation to stays within on both datasets but with larger residual scatter, delimiting the law's cleanest domain to the mid-noise range . Memory, we have shown, is an invertible, stackable layer, not an uninterpretable black-box capacity: its cost to generalization obeys a quantitative causal law with a calibrated constant that clusters near across datasets and class counts at the shared reference capacity, with its capacity dependence mechanistically traced.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.