Spectral Suppression for Annotation-Free Bias Mitigation
Abstract
Deep neural networks trained with Empirical Risk Minimization (ERM) often rely on spurious correlations rather than robust features, particularly when mul- tiple bias attributes are strongly correlated with the target label. Existing debias- ing methods typically require expensive per-sample bias annotations, which limits their practicality in real-world settings. We introduce Spectral Alignment Sup- pression (SAS), a novel annotation-free debiasing method. SAS identifies the dominant spectral subspace of the centered feature matrix and attenuates classifier feature coordinates with high leverage in that subspace. This is combined with a cosine classifier that makes the logits invariant to uniform feature scaling and per- class prototype scaling. Across Biased-MNIST, CelebA, and MS-COCO, SAS improves substantially over ERM, surpasses annotation-free two-step baselines (JTT, LfF) in the majority of evaluated settings, and remains competitive with su- pervised debiasing methods without using bias annotations. A balanced spectral audit shows dataset-dependent geometry: bias dominates the leading subspace on Biased-MNIST, co-locates with target information on CelebA, and is strongly de- codable together with the target in a compact top-32 subspace on MS-COCO. On unbiased benchmarks (CIFAR-10, STL-10, and MNIST), SAS maintains accuracy comparable to ERM.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.