Asymmetrically Unfair Decision Making and Reasoning in LLMs
Abstract
As LLMs are being deployed in higher-stakes tasks such as hiring, healthcare, and law, developers are making efforts to align them such that their decision making and processes are fair and unbiased, to avoid harms and discrimination. Yet, existing benchmarks only measure fairness at a surface level, risking missing subtle discrimination. To address this, we introduce the FairStress benchmark, which includes 300 decision-making cases across 29 critical domains (e.g., job candidate selection) with 4 degrees of demographic identity signaling, from implicit to explicit. In contrast to previous fairness benchmarks, FairStress examines LLM decisions as a full process, including output-level, (a)symmetric robustness to adversarial prefixes that subtly push towards certain identities, and verbalized and internal reasoning. Contrary to the stereotypical LLM behavior reported by prior work, comprehensive benchmarking with \FairStress shows anti-stereotypical overcorrection, an asymmetry that emerges more strongly in how robustly the model makes decisions for one person over another. With explicit identity, accuracy differs by only 2.9 points between groups, yet one added sentence makes 15 models abandon correct answers favoring the majority 9.3 points more often, and flip ties toward the minority 23.9 points more often. To test mitigations to address this, we use 9 post-hoc debiasing methods and show that they do not reliably reduce the asymmetry. We further trace the behavior back to the training data and find asymmetric demographic framing, with minorities portrayed and protected more positively. Therefore, we argue that fairness evaluations should test both the direction and robustness of a decision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.