acceptodds
Under review as a conference paper at ICLR 2027

When Observational Context Overrides Interventional Evidence in Transformers

Abstract

Can a transformer use interventional evidence when its context also contains conflicting observations? We study this question in synthetic structural causal models where observational association and the interventional effect have opposite signs, training decoder-only transformers from scratch on symbolic records while varying the training mixture and the context composition separately. We first formalize what each record type identifies: interventional probes determine the queried effect, observational records do not, and an ideal reader can still be harmed by observations on worlds where the two disagree. Against this reference, the composition of the context strongly shapes the answer. On 50 positive-effect evaluation worlds, a probe-only model produces 41 correct-sign slopes and 4 reversals, compared with 18 and 19 for a separately trained mixed-context model. At a single fixed checkpoint, adding observational records while keeping four probes reduces correct-sign slopes from 31 to 17 on 37 commonly parsed worlds and raises reversals from 1 to 11, an outcome that an ideal reader under the positive-effect training prior avoids, and replacing the observational values with zeros yields 25 correct-sign slopes of 33, compared with 16 for the world's own values. Activation patching shows that representations at the observational-record positions are sufficient to carry much of this effect in the tested subsets. The models also exhibit a positive-sign prior: on binary causal graphs, both models reverse on most negative-effect worlds, and sign-balanced retraining improves in-distribution sign reading for the probe-only model while leaving the out-of-distribution positive offset in place. Evaluations of causal learning should therefore vary inference-time evidence composition, not only the training mixture.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.