acceptodds
Under review as a conference paper at ICLR 2027

Diagnosing and Mitigating Late-Stage Shortcut Emergence in Generative Self-Supervised Learning

Abstract

Generative self-supervised learning is commonly improved by extending pretraining, yet better in-distribution representations do not necessarily translate into uniformly better robustness. We study this discrepancy by tracking frozen representations throughout prolonged masked-image pretraining. Beyond controlled stylized shifts, we find that performance on several style-sensitive natural distribution shifts, including ImageNet-Sketch and a label-aligned subset of ImageNet-R, can peak substantially earlier than in-distribution accuracy and subsequently decline while in-distribution performance continues to improve. This behavior recurs across multiple settings but is shift-dependent, revealing a late-stage divergence between in-distribution fitting and transferable representation quality. Motivated by this observation, we formulate an operational invariance-aligned pretraining design principle: nuisance variation should not only be introduced at the input, but should also be paired with supervision that makes nuisance-dependent features unreliable for minimizing the pretraining objective. We instantiate this idea with HyGDL, which combines cross-perturbation self-distillation, low-rank geometric projection, and residual-conditioned reconstruction to encourage perturbation-stable representations. Across unsupervised domain generalization benchmarks, HyGDL improves robustness over generative SSL baselines and mitigates the late-stage degradation observed in standard pretraining. Representation-level probing further shows that the learned stable projection preserves substantially more semantic predictability than its orthogonal residual, while the residual remains non-trivially domain-predictive. Together, these results characterize a previously underexplored failure mode of prolonged generative pretraining and provide an operational approach for reducing shortcut-sensitive representation drift.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.