acceptodds
Under review as a conference paper at ICLR 2027

Disc-InfoMax: Self-Supervised Learning of Invariant Causal Representations

Abstract

InfoMax self-supervised learning (SSL) has driven significant advances in representation learning, but standard approaches often entangle true causal variables with spurious correlations, leading to catastrophic failures under distribution shifts. While multi-view causal representation learning offers theoretical guarantees for identifying true causal variables, practical implementations in SSL settings are bottlenecked by their reliance on explicit view labels and highly non-convex optimization. To bridge this gap, we address the broader challenge of causal representation learning without explicit view labels, which includes the InfoMax SSL setting. We make two primary contributions. First, we introduce a novel training objective, Disc-InfoMax, which dynamically infers a shared content mask from the empirical differences between paired representations. We prove that under asymptotic convergence, the optimal encoder is theoretically guaranteed to block-identify true underlying causal variables. Second, we propose BayesDisc as a practical algorithm that effectively mitigates non-convexity, only requiring standard deep learning optimizers. Finally, we demonstrate on several datasets that BayesDisc successfully improves causal variable identification over InfoMax SSL baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.