Diagnosing and Mitigating Late-Stage Shortcut Emergence in Generative Self-Supervised Learning
Abstract
Generative self-supervised learning is commonly improved by extending pretraining, yet better in-distribution representations do not necessarily translate into uniformly better robustness. We study this discrepancy by tracking frozen representations throughout prolonged masked-image pretraining. Beyond controlled stylized shifts, we find that performance on several style-sensitive natural distribution shifts, including ImageNet-Sketch and a label-aligned subset of ImageNet-R, can peak substantially earlier than in-distribution accuracy and subsequently decline while in-distribution performance continues to improve. This behavior recurs across multiple settings but is shift-dependent, revealing a late-stage divergence between in-distribution fitting and transferable representation quality. Motivated by this observation, we formulate an operational invariance-aligned pretraining design principle: nuisance variation should not only be introduced at the input, but should also be paired with supervision that makes nuisance-dependent features unreliable for minimizing the pretraining objective. We instantiate this idea with HyGDL, which combines cross-perturbation self-distillation, low-rank geometric projection, and residual-conditioned reconstruction to encourage perturbation-stable representations. Across unsupervised domain generalization benchmarks, HyGDL improves robustness over generative SSL baselines and mitigates the late-stage degradation observed in standard pretraining. Representation-level probing further shows that the learned stable projection preserves substantially more semantic predictability than its orthogonal residual, while the residual remains non-trivially domain-predictive. Together, these results characterize a previously underexplored failure mode of prolonged generative pretraining and provide an operational approach for reducing shortcut-sensitive representation drift.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.