acceptodds
Under review as a conference paper at ICLR 2027

Evaluating Weight-Level Bias Mitigation Methods Against Behavioral and Activation-Space Bias

Abstract

Ensuring language models are free of harmful bias matters wherever their outputs shape real decisions. Yet, bias evaluation has relied almost entirely on behavioral metrics, leaving open whether mitigation methods actually remove bias from a model's internal representations. Fewer studies still examine whether weight-level methods, techniques that directly manipulate a model's weights, can eliminate both behavioral and activation-space bias simultaneously. We test whether ablation, low-rank adaptation (LoRA), and orthogonal fine-tuning (OFT) can remove injected behavioral and activation-space bias across different LLM families and model sizes. We devise a novel double-contrast, mean-of-differences bias identification technique that selects the layer where the bias signal in the model's representations is strongest and most consistent with the contrasting demographics. Our experiments show that injected bias becomes more concentrated as layer depth increases, but this correlation does not establish which layers cause the bias. Ablation fails to remove behavioral bias and does not restore the model's task capability, even when constructed explicitly to remove the bias direction. LoRA and OFT remove behavioral bias in all cases, but their effect on activation-space bias varies across models. These findings show that evaluating behavioral or activation-space bias alone is insufficient, motivating the need to assess both simultaneously.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.