acceptodds
Under review as a conference paper at ICLR 2027

Relational Steering with a Fixed Direction: Causal Effects and Behavioral Limits

Abstract

Can a single fixed activation direction produce opposite but appropriate effects across contexts? We study this question with a factorial design crossing epistemic alignment or distortion with favorable or unfavorable outcomes. In Qwen3-8B-Base and Llama-3.1-8B, we construct a relational direction from the factorial interaction contrast, which cancels the designed additive main effects, freeze it before held-out evaluation, and test causal writability through teacher-forced consequence preferences. The aggregate relational effect transfers to independent novel domains in both model families, although strict branchwise transfer is established only in Llama. Structured controls show that response magnitude alone does not explain the effect, and a discovery-selected best-of-1,024 Gaussian direction shows the branchwise sign pattern is not unique to the factorial direction, which nevertheless has a substantially larger held-out local response. Controlled-readout writability does not translate monotonically into open-ended behavior: under prompt-prefill-only intervention in Llama, the distortion branch is stronger under teacher forcing, whereas the alignment branch has the larger and more persistent trajectory-average behavioral effect, a pattern reproduced by an independent automated audit and, for the average effect, by blinded human calibration. Relation controls further constrain interpretation. Removing the stated causal linkage weakens steering in both model families, whereas explicit reversal does not produce the opposite joint pattern; a post-hoc baseline-conditioned diagnostic shows the Llama reversal response retains a default-oriented component not explained by simple amplification of the contemporaneous preference, while Qwen is substantially more baseline-tracking under the same manipulation. These results establish context-responsive relational writability while delimiting what it implies: the observed sign structure is not unique to the factorial direction, controlled writability does not guarantee persistent behavioral expression, and relation sensitivity does not by itself establish mechanism-sensitive causal-rule recomputation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.