acceptodds
Under review as a conference paper at ICLR 2027

Efficient and Interpretable Transformer for Counterfactual Fairness

Abstract

The growing reliance on machine learning for decision-making in highly regulated industries has raised fundamental challenges in regulatory fairness and interpretability compliance. Although attention-based transformers provide a powerful framework for modeling complex feature dependencies, existing debiasing methods for tabular data either fail to guarantee counterfactual fairness or rely on restrictive causal assumptions that limit practical applicability. We propose Feature Correlation Transformer (FCorrTransformer), an efficient and attention-light architecture specifically designed for tabular learning, where the attention matrix offers direct interpretation as pairwise feature dependencies. Building on this structure, we introduce Counterfactual Attention Regularization (CAR), a regularization framework that enforces group-invariant representations of sensitive features at the attention level, thereby yielding counterfactually fair predictions. We further develop an optimized input augmentation strategy for efficient counterfactual attention learning with minimum computation resource. Empirical experiments on both imbalanced classification and regression benchmarks demonstrate that FCorrTransformer+CAR achieves strong counterfactual fairness guarantees with preserved predictive performance, offering a practical framework for responsible AI in regulatory-sensitive domains.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.