I CAN Debias: Uncovering and Enhancing Self-debiasing Mechanisms of Social Bias in LLMs
Abstract
As LLMs rapidly advance and become increasingly agentic, mitigating social bias and ensuring fairness at scale have become critical challenges. Prior work has shown that LLMs exhibit an emergent self-debiasing capability—an intrinsic mechanism that enables them to avoid generating biased outputs—but this mechanism remains poorly understood, leaving existing bias-mitigation approaches heavily reliant on carefully curated prompts or inefficient training procedures that risk catastrophic forgetting. To address this gap, we propose I-CAN-Debias, a unified framework for uncovering and enhancing self-debiasing mechanisms in LLMs through behavior-grounded Consistency–contrast-Aware Neuron analysis, which defines self-debiasing neurons not by the magnitude of their responses to biased inputs but by activation patterns that distinguish successful debiasing from biased failures. Using I-CAN-Debias, we find that self-debiasing neurons account for less than 2% of model parameters yet are causally critical: deactivating them raises the biased-response rate above 95%. These neurons cluster in the query and value projections of later layers and exhibit highly asymmetric cross-category transferability. Furthermore, targeted fine-tuning of these neurons substantially improves fairness while preserving general capabilities. Importantly, models tuned with I-CAN-Debias generalize robustly across diverse evaluations beyond the training distribution, spanning open-ended fairness, jailbreak attacks, and implicit bias, without requiring test-time debiasing prompts. Together, these findings provide new mechanistic insights and a parameter-efficient approach toward developing fair and trustworthy AI systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.