What Can the Representation Space Reveal About Memorization?
Abstract
Memorization plays an important role in how neural networks fit rare and atypical training examples, but measuring per-example memorization is computationally expensive: estimating the Feldman–Zhang memorization score requires training many models. Existing practical proxies primarily rely on losses, gradients, or loss-surface geometry. We investigate whether memorization is directly reflected in the geometry of learned representations. We propose a single-run proxy that tracks, throughout training, how strongly an example's position relative to class centroids in representation space disagrees with its observed label. Under explicit assumptions, we relate this trajectory-averaged disagreement to Feldman–Zhang memorization through a sandwich bound, linking our score to memorization weighted by the fraction of training before the sample’s representation becomes consistently aligned with its assigned class. Empirically, our score achieves the strongest agreement with Feldman–Zhang memorization among the evaluated proxies on ImageNet, while remaining competitive on CIFAR-100. The same signal is highly effective for mislabel detection and true-label recovery. For mislabel detection on CIFAR-100 under within-superclass label flips (asymmetric noise), our method achieves the highest AUROC across nearly all evaluated noise rates. Its AUROC decreases by less than one percentage point as the noise rate increases from 5% to 30%, compared with drops of up to approximately twelve percentage points for competing methods. These results suggest that memorization leaves a persistent and measurable geometric signature in learned representations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.