A-IEOF: Anchored Interventional Equalized-Odds Fairness for Feature Addition
Abstract
Machine learning systems in high-stakes domains must balance utility with group fairness. Conventional debiasing optimizes fairness on the training distribution and does not address disparities that appear under distributional shift. We introduce Anchored Interventional Equalized-Odds Fairness (A-IEOF), an in-processing framework for feature addition, where a newly available feature block B can improve utility but may cause fairness regressions. A-IEOF trains a student model with a minimax objective that (i) enforces worst-case multi-group equalized odds under adversarially generated counterfactual replacements of B, (ii) anchors predictions to a no-B teacher to preserve behavior on shared inputs, and (iii) stabilizes calibrated scores against variation in B. On a multi- institutional clinical corpus, A-IEOF attains the highest AUC (0.976) and lowest equalized-odds gap (0.085–0.126) across global, per-group, and Hardt-EO post-processing policies, with better calibration (Brier 0.0495–0.0507; ECE 0.0262–0.0343) than adversarial gradient-reversal (GRL) debiasing. To test the robustness claim directly we introduce a controlled tabular feature-addition benchmark in which P (B |S) can be shifted by a prescribed total-variation amount: A-IEOF holds its equalized-odds gap essentially flat as the shift grows (degradation < 0.01 up to ∆T V = 0.5), consistent with our transfer bound, while an empirical-risk baseline is four times more sensitive to B. Over ten seeds A-IEOF matches a constrained-optimization (reductions) baseline in distribution and surpasses invariant-learning baselines (IRM, CORAL) and GRL. On three real tabular datasets (UCI Adult, COMPAS, CT metadata) it attains the lowest equalized-odds gap on Adult and COMPAS and the lowest sensitivity to B on all three. A-IEOF thus offers a principled, auditable route to feature addition that aligns utility, calibration, and fairness under realistic deployment conditions. Our implementation and benchmark are available at https://anonymous.4open.science/r/AIEOF-778E.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.