Many Writers, One Reader: A Causal Handle on Rule Polarity in Language Models
Abstract
We show that language models represent whether a rule permits or prohibits an action in a form that causally controls the verdict. We test this across four instruction-tuned models (Llama, Qwen, Mistral, Gemma) in a controlled anti-structuring task based on the Bank Secrecy Act. At a late decision layer, a single direction separates otherwise matched rules that permit or prohibit the same action. Removing this coordinate sharply degrades the decision, while repeatedly replacing the same coordinate through depth recovers 73–93% of the opposite decision. The computation that produces this coordinate is distributed: many late-layer components write this coordinate, while a single direction provides a compact causal readout. This effect is rule-selective. Across matched control tasks, it cannot be explained by either a generic binary-decision direction or the SATISFIED/EVADED answer tokens alone. The model's own computation also moves this coordinate: injecting a false premise and the conclusion it implies into a genuine violation shifts the internal state toward the same coordinate occupied by genuine permission. Finally, we test the same causal signature in FAA drone-altitude rules. It reproduces strongly in Llama and Gemma and more weakly in Qwen; Mistral does not form a sufficiently strong FAA decision signal for the downstream causal test.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.