Direct Hidden-State Alignment: Mapping and Controlling Preference Expression in LLMs
Abstract
In many settings, post-training need not create the target behavior from scratch: the base model can already produce it, but not reliably. This shifts part of preference alignment from capability acquisition to behavioral expression. This raises a central question: if target-supporting computation is already present, what prevents the preferred behavior from reliably dominating generation? At each generation step, many Transformer components write to the same residual stream, yet their effects are combined into a single next-token distribution. A target-supporting computation can therefore be present yet be outweighed by other computations. We hypothesize that reliable expression depends on this internal competition during inference. We introduce Residual Competition Maps (RCMs), which map a specified preference onto native residual computation by measuring signed causal effects relative to that preference. Across preference domains, RCMs reveal residual competition whose prevalence varies by task, distinct component roles relative to the target preference, and cases where a single native-component intervention reverses the preference outcome. DPO substantially reorganizes these effects and can weaken opposing effects, which may nevertheless persist. Beyond analysis, RCM provides causal guidance for where to intervene. To directly control preference formation during inference, we propose Direct Hidden-State Alignment (DHSA), which treats inference-time hidden states rather than base-model weights as the direct adaptation space. RCM-guided Causal Activation State Transition (CAST) instantiates this principle through local state interventions at a small number of preference-relevant interfaces while freezing the base model. With only 256–16,384 controller parameters, CAST reaches DPO-competitive operating points across three preference domains, can complement DPO-trained models, and can be enabled or removed at inference time.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.