acceptodds
Under review as a conference paper at ICLR 2027

RL-SAF : Reinforcement Learning Based Semantic Adversarial Fairness

Abstract

Adversarially trained models can achieve strong overall robustness while exhibiting substantial class-wise performance disparities under attack. Although this adversarial fairness problem has been studied in image classification, it remains largely unexplored in semantic segmentation. We show that such disparities occur across datasets, architectures, and adversarial training strategies, that adversarial perturbations amplify rather than uniformly degrade class-wise performance, and that class frequency alone does not explain which classes are most vulnerable. We further show that adapting a state-of-the-art adversarial fairness method from image classification to semantic segmentation reduces both worst-class and overall robustness, which we attribute to a mismatch between recall-based notions of class difficulty and the IoU objective of dense prediction. Motivated by these findings, we propose RL-SAF (Reinforcement Learning-Based Semantic Adversarial Fairness), which combines class-level loss aggregation with a policy-gradient controller that adaptively learns class-wise loss weights from validation feedback. On Cityscapes with DeepLabV3, RL-SAF improves the adversarial IoU of the most vulnerable classes across all evaluated attacks and preserves robust mIoU under the attack used for training, while trading off part of it under stronger attacks. These gains come at a substantial cost in clean accuracy, revealing a clear trade-off between adversarial fairness and standard segmentation accuracy. Overall, our results establish class-wise adversarial robustness disparities as an important challenge in semantic segmentation and provide initial evidence to motivate adaptive, fairness-aware adversarial training for dense prediction tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.