acceptodds
Under review as a conference paper at ICLR 2027

I CAN Debias: Uncovering and Enhancing Self-debiasing Mechanisms of Social Bias in LLMs

Abstract

As LLMs rapidly advance and become increasingly agentic, mitigating social bias and ensuring fairness at scale have become critical challenges. Prior work has shown that LLMs exhibit an emergent self-debiasing capability—an intrinsic mechanism that enables them to avoid generating biased outputs—but this mechanism remains poorly understood, leaving existing bias-mitigation approaches heavily reliant on carefully curated prompts or inefficient training procedures that risk catastrophic forgetting. To address this gap, we propose I-CAN-Debias, a unified framework for uncovering and enhancing self-debiasing mechanisms in LLMs through behavior-grounded Consistency–contrast-Aware Neuron analysis, which defines self-debiasing neurons not by the magnitude of their responses to biased inputs but by activation patterns that distinguish successful debiasing from biased failures. Using I-CAN-Debias, we find that self-debiasing neurons account for less than 2% of model parameters yet are causally critical: deactivating them raises the biased-response rate above 95%. These neurons cluster in the query and value projections of later layers and exhibit highly asymmetric cross-category transferability. Furthermore, targeted fine-tuning of these neurons substantially improves fairness while preserving general capabilities. Importantly, models tuned with I-CAN-Debias generalize robustly across diverse evaluations beyond the training distribution, spanning open-ended fairness, jailbreak attacks, and implicit bias, without requiring test-time debiasing prompts. Together, these findings provide new mechanistic insights and a parameter-efficient approach toward developing fair and trustworthy AI systems.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.