To Compare, or Not to Compare: On Methodological Practices in Evaluating Social Bias
Abstract
As Large Language Models are increasingly deployed in critical applications, robustly evaluating their social biases is paramount. However, the current literature suffers from widespread methodological fragmentation, which yields contradictory conclusions. As we show, this is in part related to the structural framing of benchmark-level evaluations: whether demographic targets are assessed in isolation or compared directly in a forced choice. Our evaluation across multiple model families reveals a massive, systematic paradigm gap: while isolated assessments limit prejudice activation, comparative settings act as aggressive catalysts for latent stereotypes and discrimination, a shift primarily driven by underspecified contexts. To isolate this effect, we introduce a unified and controllable framework that standardizes heterogeneous benchmarks across both paradigms and disentangles the confounding effects of Chain-of-Thought reasoning, neutral fallback options, and other structural artifacts in social bias evaluations. Alarmingly, CoT reasoning strongly exacerbates social biases under comparative settings while having virtually no effect in isolated ones, and this systemic bias persists even when models are provided neutral fallback options or claim to answer randomly. Finally, we demonstrate that this comparative prejudice is a generalized phenomenon that scales positively with model size. Ultimately, we offer a practical methodological guideline: while researchers must leverage comparative settings to robustly audit hidden biases, practitioners cannot safely rely on comparative deployments in ambiguous real-world tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.