acceptodds
Under review as a conference paper at ICLR 2027

Multigroup Social Debiasing of LLMs via Neuron Discovery

Abstract

Mitigating social biases in language models is a prerequisite for their responsible deployment. Many benchmarks and debiasing methods operationalize bias as a difference between two groups, a target group and a single alternative, so a model can reduce this pairwise gap while its preferences over the remaining groups stay biased. We define social bias as the divergence from uniform of the distribution a model induces over a set of demographic groups, with the completion held fixed so that any preference is attributable to the group. To this end, we introduce , an extension of in which each target group is complemented by distractor groups selected for maximal dispersion over real-world taxonomies, while stereotypical and anti-stereotypical completions are held fixed. We then propose DAID, a two-stage method that first identifies the MLP neurons whose suppression moves the group distribution toward uniform and then fine-tunes only these neurons' parameters. Across four models on English text, DAID reduces the median group-probability ratio from - to - on *nationality* and from - to below on *occupation*, and increases ICAT by – points, at a cost of - points on . Model-editing and activation-intervention baselines raise ICAT by up to points but leave the ratio on *nationality* between and . We find that suppressing the identified neurons without training also yields near-uniform distributions but degrades language modelling capabilities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.