What a Steered Lie Does Inside: The Deception-Belief Entanglement in LLMs
Abstract
Activation probes can monitor AI systems by decoding a model's internal belief and flagging public outputs that depart from it. This use requires the decoded belief to stay reliable when the model deceives. We test this requirement in One-Night Ultimate Werewolf, where night-phase card swaps give each seat an observation-grounded believed role that can differ from its true role. A belief probe decodes the model's believed role above 0.84 accuracy. We induce deception in Llama-3.1-8B-Instruct with activation steering added only at public-speech token positions, leaving the generation of private beliefs unchanged. The belief readout shifts monotonically with steering strength, while the model’s self-reported role becomes less confident. Removing the binary Werewolf-vs-Village direction explains only 7-10% of the shift; 83% instead concentrates in three high-gain role axes (Villager, Werewolf, and Tanner), and the steering delta is enriched 2-10 times in these directions relative to a random vector. Perturbations confined to the belief subspace flip the probe in 86-90% of cases, versus 2-7% for random subspaces. Qwen2.5-7B-Instruct shows the same structure, suggesting that belief monitors should be stress-tested against interventions on adjacent features.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.