Utility-Preserving Relabeling Reveals Semantic Sensitivity in Language-Model Decision Making
Abstract
A decision should not change when an action is given a different label but its payoff stays the same. We study whether language models satisfy this form of strategic invariance. We introduce CSS-Bench, a benchmark of Prisoner’s Dilemma, Chicken, and Stag Hunt games in which action labels are swapped while the underlying payoff structure remains fixed. Across nine open-weight models, relabeling changes decision accuracy substantially, with absolute gaps of up to 83.4 percentage points across model-vocabulary cells. The effect varies in both magnitude and direction across models and vocabularies, and is not restricted to semantically loaded labels. We then trace the effect inside the models. Label-related information becomes linearly decodable at an average normalized depth of 0.25, while activation patching identifies causal influence at a substantially later depth of 0.76 across 36 model-vocabulary cells. Thus, information about action labels can be represented well before it becomes causally relevant to the decision. Causal steering further changes decision behavior, but transfers substantially less reliably to unseen labels than the underlying representation. Together, these results show that language models can change their decisions when action labels are swapped without changing the underlying payoffs, and provide a controlled framework for studying how linguistic information influences decision-making.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.