Dual Memory Retrieval: Bridging Contrastive Learning With and Without Negative Pairs
Abstract
Contrastive learning builds self-supervision by comparing positive and negative samples. Methods such as SimSiam, DirectPred and DINO remove the negatives and resist collapse with a predictor or prototypes. Their contrast is carried jointly by the objective and these modules, rather than by the objective alone. We treat the negatives, the predictor and the prototypes alike as a memory of the representation distribution and recast contrast as retrieval from this memory, placing methods with and without negatives on one comparable axis. Revisiting contrastive learning from this perspective, we propose Dual Memory Retrieval (DMR). Its gradient is the difference between the retrievals of the two positive views at two temperatures in a softmax family or two exponents in a spectral family. We examine the resulting learning behavior on four image datasets. With a memory derived directly from the batch, the gap between the two temperatures or exponents is essential for training. With a memory learned by gradient steps, training fails when the positive enters the loss directly. Retrieving the positive through the memory lets the spectral family train with a learnable matrix, but not the softmax family without further mechanisms.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.