Causal Consequence Invariance: Do LLM Causal Predictions Depend on Downstream Preferences?
Abstract
Large language models are increasingly used to produce predictions and analyses that may influence downstream decisions. In these scenarios, the user requesting a prediction may prefer one outcome over another. If such preferences change the model's estimate while the underlying causal evidence remains unchanged, then the prediction is less trustworthy since it no longer reflects the evidence alone. We construct 48 controlled structural causal model scenarios with exact interventional ground truth spanning confounding, collider bias, mediation, and no-adjustment settings. We find that across Claude Sonnet 5 and GPT-5, the mean absolute difference between opposing mild-preference conditions is only 0.053 and 0.004 percentage points, respectively. We find no detectable systematic variation in preference sensitivity across causal structures, and stronger stated incentives do not produce a practically meaningful monotonic response. Preference-induced variation is also comparable in magnitude to variation produced by causally irrelevant context. We further test whether this stability persists when the causal evidence becomes less precise and when abstract variables are replaced with semantically meaningful framings. Lower evidence precision produces a statistically detectable but practically negligible increase in absolute preference sensitivity for GPT-5, without a directional shift toward the preferred outcome; Claude shows no detectable increase, and semantic framing produces no detectable increase for either model. These results suggest that, in this controlled benchmark, both models' causal estimates are behaviorally invariant to the tested downstream preferences. More broadly, this work provides a practical way to evaluate preference sensitivity in the causal predictions of future models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.