Walking Across Images: Random Walk Pairing for Self-Supervised Learning
Abstract
The ability to capture both local invariances and global semantic structures is essential for effective self-supervised learning. In this paper, we address a fundamental challenge of ensuring robust generalization to unseen data while effectively capturing global semantic structures. Motivated by a spectral embedding perspective, we identify potential sources of suboptimal generalization in SSL methods that can be viewed through a spectral lens. We propose a simple yet effective training strategy that can be easily applied to a variety of SSL methods. We introduce SAG-VICReg (Stable and Generalizable VICReg) as a detailed instantiation, demonstrating improvements in generalization and global semantic understanding. Through comprehensive experiments across diverse SSL frameworks including SimCLR and TCL, we show broad applicability with consistent improvements. The same pairing also improves the fine-tuning of today's strongest pretrained backbones, with gains that persist under linear probing, grow as labels become scarce, and transfer to unseen datasets. Our enhanced methods exhibit superior performance on metrics designed to evaluate global semantic understanding while maintaining competitive results on local evaluation metrics. We also propose a label-free evaluation metric that accounts for global data structure without requiring labels–a key advantage when labeled data is scarce or unavailable.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.