Can LLMs See Through the Mind Games? Benchmarking Theory of Mind in Psychological Manipulation
Abstract
Theory of Mind (ToM) is a central ability for socially competent large language models (LLMs). Yet most existing evaluations consider benign settings, overlooking hidden intentions and adversarial social influence. Psychological manipulation poses a distinct challenge: a manipulator strategically shapes a target’s mental states while concealing the goal behind that influence. We propose ManipToM, a ToM benchmark in psychological manipulation. It tests whether LLMs can reason over asymmetric and evolving mental states, particularly distinguishing what the manipulator intends and wants from what the target believes about it. Grounded in psychology, ManipToM covers 16 manipulation mechanisms across three domains: cognitive, affective, and identity. Our controlled generation pipeline instantiates manipulative interactions, explicitly tracks agents’ evolving mental states, and uses this structured interaction state to construct and validate benchmark questions automatically. ManipToM contains 645 dialogues and 13,734 questions spanning explicit, dynamic, and applied ToM. Evaluating nine LLMs, we find that strong performance on explicit mental-state questions does not consistently extend to tracking state changes or predicting behavior, and chain-of-thought prompting does not close this gap. Our results establish manipulation as a challenging testbed for multi-agent mental-state reasoning and provide a scalable framework for evaluating models under strategic social influence. The dataset and code will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.