acceptodds
Under review as a conference paper at ICLR 2027

Benchmarking and Evaluating Implicit Biases in Single- and Multi-Agent Language Models

Abstract

As large language models (LLMs) are increasingly used for decision support, planning and multi-agent societal simulation, evaluating bias in isolated outputs is not sufficient. We also need to measure how bias signals implicitly affect model decisions. Existing bias benchmarks primarily focus explicitly on single-agent settings, leaving limited tools for studying whether multi-agent settings may mitigate, preserve, or amplify such signals. We introduce SIFT, the Surfacing Implicit bias in Foundation models Toolkit, a unified benchmark and evaluation framework for measuring implicit stereotype signals in both single- and multi-agent LLMs. SIFT-Bench contains theory-grounded decision-making scenarios spanning five task families, eighteen domains, eleven social axes and six evidence conditions, yielding 26.4K queries. SIFT-Eval evaluates models across persona and interaction conditions using metrics for stereotype rate, counterfactual and style inconsistency, demographic reasoning and interaction shifts. Across model families, we find that persona conditioning can sharply increase stereotype-consistent decisions. Additionally, multi-agent settings show +12.4 percentage points higher stereotype rates than matched single-agent evaluations in participant-centered tasks like collaborative planning and task allocation. However, initial-to-final analyses reveal both amplification and attenuation depending on the interaction condition and domain. Together, SIFT enables controlled assessment of implicit bias in model decisions, its evolution through interaction, and agentic systems before deployment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.