Masking Spurious Concepts via Sparse Autoencoders for Efficient Bias Mitigation
Abstract
Deep neural networks trained via Empirical Risk Minimisation (ERM) are prone to relying on spurious correlations: confounding features such as backgrounds that are statistically associated with a class label during training but are not causally relevant. When these correlations do not hold at test time, performance degrades on subgroups without spurious features. Existing mitigation methods either (a) require model retraining or large group label annotations, or (b) modify only the final classifier head, leaving spurious information encoded in the latent space and thus exhibiting a trade-off between worst-group and average accuracy. To circumvent these limitations, we propose SPOCK, a bias mitigation method that uses a Sparse Autoencoder to decompose latent representations into sparse concepts and learns a mask that suppresses spurious concepts while preserving those useful for the downstream task. By filtering concepts before reconstructing the representations, SPOCK reduces the model's reliance on spurious correlations while keeping the existing classifier fixed. In this way, SPOCK is compute- and label-efficient, requiring only a handful of group-annotated validation examples and avoiding end-to-end retraining. Our results across four synthetic and real-world datasets show that SPOCK substantially improves worst-group accuracy over ERM while maintaining high average accuracy, offering a more favourable trade-off than competing bias mitigation approaches.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.