Subconscious Steering
Abstract
Large language models can be steered by seemingly harmless text. We introduce subconscious steering: models can be steered by text just by the act of reading it, despite stating they will not; warnings do not fully prevent it. To measure it we define fractional authority: rather than sometimes working and sometimes failing, injected text consistently shifts a model a fixed fraction of the way toward the attacker’s objective. For shopping agents across seven models, three sentences of seemingly innocent text obtain an effect equivalent to a 20% price discount. The text need not look like a command to work. Mechanistically, we find that models simulate goals by default, even when the goal is quoted, unauthorized, or irrelevant to the task. How much that simulated value counts is a causally separate variable, and the two together predict the fractional effect.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.