OmniFair: A Unified Fairness Benchmark Across Tasks and Modalities
Abstract
Fairness evaluation of multimodal large language models (MLLMs) remains fragmented across tasks and modalities. Existing instance-aligned multi-task benchmarks remain text-only, whereas instance-aligned multimodal benchmarks remain confined to individual tasks, preventing controlled separation of task and modality effects. We introduce the Bias Semantic Unit (BSU), a five-dimensional representation that separates stereotype semantics from task- and modality-specific realizations. From 143 fairness benchmarks and 11M multilingual news articles, we construct BiasAtlas with 416K BSUs. Projecting 3,075 representative BSUs across five task paradigms and three modalities yields OmniFair, a benchmark of 46,125 instances with joint cross-task and cross-modal semantic-source alignment; human validation finds 97.9% semantic fidelity and 95.5% multimodal media effectiveness. Across 15 state-of-the-art MLLMs, instance-level discrimination decisions show almost no agreement across tasks () and limited agreement across modalities (). Category-level bias-rate rankings are likewise nearly uncorrelated across tasks (mean pairwise across semantic axes) but more consistent across modalities (). This contrast shows that task paradigms reshape the semantic risk landscape, whereas modalities alter which individual stereotypes elicit discrimination while preserving more of the category-level structure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.