acceptodds
Under review as a conference paper at ICLR 2027

Lossy Representation Compression for Out-of-Distribution Generalization

Abstract

Out-of-distribution (OOD) generalization aims to maintain predictive performance under distribution shifts. Despite extensive research, invariant learning methods often fail to deliver reliable improvements across diverse settings. A key limitation is that prediction invariance across training environments encourages the learning of causal features but does not remove non-causal features from learned representations. We first show that causal and non-causal features can occupy different directions in the representation space, such that encoding both produces more complex representations than encoding causal features alone. Based on this observation, we introduce Lossy Causal Feature Learning (LoCaL), a simple and effective framework that regularizes learned representations using the well-established lossy coding rate from information theory. Geometrically, the lossy coding rate measures the volume occupied by a representation at a given precision. Constraining this quantity encourages a lower-dimensional, more compact representation, thereby suppressing non-causal features. Extensive experiments across diverse datasets demonstrate that LoCaL outperforms a broad range of strong baselines, with most improvements being statistically significant.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.