acceptodds
Under review as a conference paper at ICLR 2027

Beyond Demographics: Counterfactual VQA for Auditing Fairness in Black-Box LVLMs

Abstract

Recent advances in large vision–language models (LVLMs) have amplified concerns about fairness, yet existing evaluations remain largely centered on demographic attributes (e.g., gender and race) and often conflate fairness with refusal behavior. In this study, we reveal an overlooked risk for LVLM fairness: decision stability degrades substantially under visual changes (e.g., social behaviors and aesthetic elements), even in proprietary systems. To audit this risk, we introduce ABCDE-Bench, a paired-image VQA benchmark that tests across five attribute categories: Aesthetics, Behavior, Culture, Demography, and Environment. We further propose refusal-aware metrics for black-box evaluation that disentangle decision instability from refusal asymmetry. Experiments show that across diverse LVLMs, counterfactual inconsistencies are widespread, and LVLMs consistently show lower stability under non-demographic than demographic changes, whereas human performance shows the opposite tendency. Finally, we examine lightweight mitigation strategies and find that human-norm exemplars increase consistency more uniformly across models than instruction prompting alone. Refusal-aware analysis further shows that the observed consistency gains under these interventions are accounted for by increased symmetric refusal, while the contribution from matching non-refusal answers decreases. Overall, our findings highlight an underexplored non-demographic source of instability and motivate broader, refusal-aware audits of LVLM fairness beyond demographics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.