acceptodds
Under review as a conference paper at ICLR 2027

InvariFair: One Regularization Parameter Jointly Bounds OOD Risk and Equalized Odds

Abstract

Out-of-distribution (OOD) robustness and algorithmic fairness are typically treated as separate problems that need separate machinery: invariant learning methods deal with spurious correlations between inputs and labels, while fairness methods deal with spurious correlations between protected attributes and outcomes. We observe that under a latent structural causal model (SCM), these are the same problem viewed from two angles; both happen when a predictor leans on a latent factor that is not causally related to the label. We introduce a causal-disentanglement framework for OOD generalization that provably bounds worst-case risk under a bounded shift in the spurious latent's distribution, and show that allowing the protected attribute to depend on both the causal and spurious latents instead of the spurious latent alone leaves a purely causal classifier generally with a nonzero residual demographic-parity (DP) gap — and prove that this residual gap is provably never smaller than the corresponding equalized-odds (EO) gap, with the ratio between them controlled by a single interpretable quantity: how much label noise stays once the causal factors are known. We further introduce a lightweight fairness-aware disentanglement objective that explicitly targets the spurious component of the protected attribute by routing it into the spurious latent space during training. Like all methods compared here, ours requires no protected attribute access at inference time. On two real-world fairness benchmarks, evaluated under an equal training budget, with independently-reseeded runs, and with 10 seeds per method to support paired significance testing, our method reduces the equalized-odds gap by 30-34% relative to standard Empirical Risk Minimization (ERM) and Adversarial Debiasing on Adult Income (p<0.005 for both). On COMPAS, we do not find a statistically robust improvement: our method's equalized-odds and demographic-parity gaps are not significantly different from ERM's or Adversarial Debiasing's, and are directionally (but not significantly) worse than Reweighting's. This is a weaker result than an earlier and under-powered 3-seed run of the same experiment showed, and explains why we report 10-seed statistics rather than 3. An ablation study and a real-data sweep over the invariance parameter reveal a dataset-dependent pattern: on Adult, the fairness-aware heads alone account for nearly the full improvement; on COMPAS, the heads and the invariance penalty are complementary within our method's own variants, and increasing the invariance parameter creates a further fairness-gap reduction consistent with our theoretical bound, even though this does not go into an advantage over the external baselines at n=10. Consistent with our theory, our causal mechanism improves equalized odds substantially on Adult while leaving its demographic-parity gap statistically unchanged from ERM, and we connect this same asymmetry to a corresponding in-distribution accuracy cost on Adult that the mechanism's adversarial component predicts but does not fully explain on its own.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.