How Conversations About Values Influence Choices in Language Models
Abstract
Malleability of language model responses to values expressed by users can be a two-edged sword. It allows alignment with user preferences, but can also allow influence to extend beyond the issue currently being discussed. We investigate whether values expressed by a user carry influence into choices the model makes later, in separately written situations. In scripted conversations, users argue for one side of a value contrast for six turns, then engage in 0,1, or 3 neutral turns, ending with forced-choice dilemmas involving the same value contrast. Across Gemma-2-27B, Qwen3-32B, and Llama-3.3-70B, we find that in issues of conservation versus openness to change, the models shift later choices. In contrast, in issues concerning moral values of self-transcendence versus self-enhancement, model shift was minor. The effect on openness to change versus conservation values was robust to number of neutral turns, but decreased in face of variability in content of value argument by user. In matched experiments on Gemma-2-27B and Qwen3-32B, the same value position produced a larger downstream shift when presented as the user's own rather than attributed to a colleague. Only direct activation steering on Gemma-2-27B, shifted self-transcendence versus self-enhancement choices. Overall, our results indicate that conversational influence can extend beyond immediate agreement to later choices, while its measured strength varies across statement sets and attribution framing. However, this effect did not extend to moral values, that changed only in face of direct intervention.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.