Disentangling Regularization Effects in Self-Supervised Learning through Controlled Gradient Interventions
Abstract
Multi-objective self-supervised visual learning couples auxiliary losses and geometric regularizers through shared encoders, yet regularization gains need not translate into gains over a base-loss reference. Existing studies document configuration dependence, but removing an auxiliary loss changes updates to several parameter groups, obscuring the contribution of direct encoder updates. We therefore study how direct auxiliary encoder updates influence regularization effects. At the same computational state, we retain or remove only the auxiliary encoder gradient while matching forward losses and all other parameter gradients. Crossing this intervention with regularizer inclusion, alongside a base-loss reference, distinguishes regularization effects, combination effects, and update-path dependence. Paired experiments cover controlled-background predictive learning and natural-image self-distillation across regularizers, auxiliary weights, and training stages. Results show that removing auxiliary encoder gradients reverses the mean effect of Gram regularization on accuracy from +1.98 to −1.70 percentage points (pp) in five-fold linear evaluation of Caltech–UCSD Birds-200-2011 pretraining images after 20 epochs, consistently across six runs. This reversal also occurs with covariance regularization and with Gram at a lower auxiliary weight; these combinations remain below the base-loss reference. On Tiny ImageNet after 100 epochs, Kozachenko–Leonenko regularization lowers mean held-out accuracy by 0.50 pp with auxiliary encoder updates retained, although the full DINOv2-style combination exceeds the reference by 0.92 pp. On a fixed validation set, the interaction changes by −1.04 pp between epochs 20 and 100 (95% confidence interval: −1.84 to −0.25). These comparisons inform regularizer reuse and configuration selection under specified objectives, gradient paths, and training budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.