BRIDGE: A Bias Register-Integrated Dataset for Generalized Evaluation of LLM Bias Detection
Abstract
Large Language Models(LLMs) are increasingly being used to identify discriminatory or harmful language in institutional settings, where text is often formal, legalistic, and policy-oriented. However, existing bias evaluation frameworks for LLMs largely rely on informal, synthetic, crowd-sourced or social-media-style text (BBQ, StereoSet, CrowS-Pairs, ToxiGen, Social Bias Frames), creating a structural mismatch between benchmark registers and deployment registers. Across formal and informal registers, we observe a systematic failure mode, which we call the Register Polarity Flip (RPF): eight frontier LLMs under-detect harmful bias in formal text while over-flagging neutral informal text as harmful, with consistent directionality across models. To study this failure mode, we introduce BRIDGE (Bias Register-Integrated Dataset for Generalized Evaluation), a 19,420-sentence benchmark spanning formal institutional text and informal social media text under a unified three-way taxonomy of harmful, harmless, and antibias content assembled from seven open-source corpora spanning 150 years of discourse. We also introduce GRDC, a five-metric diagnostic suite for group-conditional bias evaluation, capturing recall divergence, directional consistency, polarity flips, confusion drift, and capability-conditional drift. The RPF pattern replicates on an independent external dataset (BABE + Measuring Hate Speech) and is not meaningfully mitigated by few-shot prompting. Because source, era, topic, and genre are correlated in the BRIDGE benchmark, we further conduct a bidirectional matched-pair control (Paired Register Effect: PRE) which shows that the same register polarity direction persists. These findings show that current bias benchmarks can misrepresent model reliability for formal-text bias detection and motivate register-aware evaluation protocols that report stratified performance across deployment-relevant registers before LLMs are used in institutional auditing workflows.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.