acceptodds
Under review as a conference paper at ICLR 2027

One Neuron, Multiple Traits: Activation-Interval Neuronal Gating for Controllable Personality in LLMs

Abstract

Controlling personality expression in Large Language Models (LLMs) supports personalized dialogue with specified behavioral styles. Existing activation-based personality control methods mainly modify entire activation vectors or selected neurons, but they will inevitably undermine general ability and non-target traits due to the polysemy and overlap of neuronal logits. To study this problem, we first conduct an empirical experiment showing that distinct personality traits and general reasoning tasks share individual neurons but occupy distinct activation intervals. Building on this observation, we propose an Activation-Interval Neuronal Gating (AING) method for personality control. This method aims to identify the activation intervals within each neuron without impairing general capabilities or other traits. Specifically, AING first selects neurons by comparing paired high- and low-trait responses of LLMs by the BIG5-CHAT dialogue dataset. Then we localize trait attribution to activation intervals within each selected neuron. These intervals define which activation values receive an intervention. During inference, AING applies adaptive scaling only within the target activation interval, with the update direction and strength determined by paired high-low differences. Extensive experiments across two model families demonstrate that AING achieves precise personality steering with high target alignment, while demonstrating remarkable selectivity by minimizing collateral shifts in non-target traits. Furthermore, our approach ensures the preservation of the model’s core functional integrity, maintaining robust performance across 6 diverse general-capability benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.