acceptodds
Under review as a conference paper at ICLR 2027

Modulating LLM Behavioral Preferences by Reweighting Just a Few Weights

Abstract

Post-training can instill dispositional behaviors in large language models (e.g., refusal & sycophancy behaviors). We find that such behavioral preferences can be removed or mitigated by removing or downweighting just a few corresponding weights, or even amplified by upweighting the few. Our proposed approach, BLADE (Behavioral Localization via Activation-Difference Estimation), is a gradient-free, forward-only method that scores weights and identifies those that contribute to behavioral preferences across layers, enabling removal or modulation of an LLM’s behavioral preferences to a desired direction. We show that, through extensive experiments on various behaviors and models, our approach effectively and efficiently modulates the model’s behavioral preferences.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.