acceptodds
Under review as a conference paper at ICLR 2027

Tracing Group Fairness through Training

Abstract

When a trained model performs much worse on some demographic groups than others, a natural question is: which training examples contribute to the disparity? Fairness-aware learning typically aims to reduce such gaps without attributing them to individual training examples, while standard data attribution methods focus on predictive performance rather than fairness. We propose FAIR-DS, which assigns each example a score for its accumulated contribution to a chosen group-aware objective in a single training run. Building on In-Run Data Shapley, the method uses group validation statistics to identify harmful examples under popular fairness criteria. Our analysis bounds the attribution error and connects the computed scores to changes in the selected fairness criteria. These guarantees clarify how individual examples contribute to fairness changes along the training trajectory. Empirical studies on label-bias detection and attribution-guided data repair demonstrate competitive performance on real-world tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.