acceptodds
Under review as a conference paper at ICLR 2027

TRAIT: selecTive and Robust Activation-space Intervention for Personality Trait Control

Abstract

Large language models are increasingly used to simulate human behavior, enabling social simulation experiments. In such experiments, researchers typically control the personality of these models through prompting, fine-tuning, or activation steering. However, existing methods still fall short in the selectivity and robustness of control, either perturbing unrelated dimensions or failing to sustain the preset level across contexts, thereby undermining experimental validity. To this end, we propose TRAIT, an automated framework for selective and robust personality control. On the extraction side, Direction Extraction from Deliberation Spans (DeliSpan) takes items from the psychometric scale of the target construct as seeds, has the target model synthesize its own decision scenarios, reason through them, and make decisions, and extracts direction vectors solely from the reasoning segments in which the model reveals its own preferences while weighing trade-offs. On the injection side, Dynamic Injection with Context-Adaptive Dosing (CtxDose) trains a dose network, supervised by the deviation between post-injection scale scores and preset scores, to predict the injection strength based on context. Evaluated on the IPIP-300 inventory within the psychometric framework of Serapio-Garc´ıa et al. (2025), TRAIT sets a new representation engineering state of the art in selectivity and robustness. These suggest personality can be controlled precisely and stably, providing a foundation for variable control in social simulation experiments.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.