acceptodds
Under review as a conference paper at ICLR 2027

Masking Spurious Concepts via Sparse Autoencoders for Efficient Bias Mitigation

Abstract

Deep neural networks trained via Empirical Risk Minimisation (ERM) are prone to relying on spurious correlations: confounding features such as backgrounds that are statistically associated with a class label during training but are not causally relevant. When these correlations do not hold at test time, performance degrades on subgroups without spurious features. Existing mitigation methods either (a) require model retraining or large group label annotations, or (b) modify only the final classifier head, leaving spurious information encoded in the latent space and thus exhibiting a trade-off between worst-group and average accuracy. To circumvent these limitations, we propose SPOCK, a bias mitigation method that uses a Sparse Autoencoder to decompose latent representations into sparse concepts and learns a mask that suppresses spurious concepts while preserving those useful for the downstream task. By filtering concepts before reconstructing the representations, SPOCK reduces the model's reliance on spurious correlations while keeping the existing classifier fixed. In this way, SPOCK is compute- and label-efficient, requiring only a handful of group-annotated validation examples and avoiding end-to-end retraining. Our results across four synthetic and real-world datasets show that SPOCK substantially improves worst-group accuracy over ERM while maintaining high average accuracy, offering a more favourable trade-off than competing bias mitigation approaches.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.