acceptodds
Under review as a conference paper at ICLR 2027

OmniFair: A Unified Fairness Benchmark Across Tasks and Modalities

Abstract

Fairness evaluation of multimodal large language models (MLLMs) remains fragmented across tasks and modalities. Existing instance-aligned multi-task benchmarks remain text-only, whereas instance-aligned multimodal benchmarks remain confined to individual tasks, preventing controlled separation of task and modality effects. We introduce the Bias Semantic Unit (BSU), a five-dimensional representation that separates stereotype semantics from task- and modality-specific realizations. From 143 fairness benchmarks and 11M multilingual news articles, we construct BiasAtlas with 416K BSUs. Projecting 3,075 representative BSUs across five task paradigms and three modalities yields OmniFair, a benchmark of 46,125 instances with joint cross-task and cross-modal semantic-source alignment; human validation finds 97.9% semantic fidelity and 95.5% multimodal media effectiveness. Across 15 state-of-the-art MLLMs, instance-level discrimination decisions show almost no agreement across tasks () and limited agreement across modalities (). Category-level bias-rate rankings are likewise nearly uncorrelated across tasks (mean pairwise across semantic axes) but more consistent across modalities (). This contrast shows that task paradigms reshape the semantic risk landscape, whereas modalities alter which individual stereotypes elicit discrimination while preserving more of the category-level structure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.