Interpreting and Steering LLM Agents for Social Simulations
Abstract
Generative agents built on large language models (LLMs) have proven powerful for simulating human behavior, making them valuable additions to the social scientific toolkit. However, the black-box nature of LLMs limits their value for simulation by constraining both: (i) interpretability: uncovering mechanisms driving observed behavior; and (ii) steering: muting or amplifying mechanisms of action along theoretically meaningful dimensions. Here, we demonstrate how the black box could be opened up to further enrich LLM-based simulations. We illustrate our approach by interpreting and steering two primitive components of human behavior, namely preferences (risk attitudes, fairness preferences) and capabilities (divergent creativity, product innovation), using four classic economic and creative tasks implemented as natural-language interactions. We demonstrate the utility of sparse autoencoder (SAE)-based techniques for interpreting mechanisms and compare three types of steering methods: (1) prompt-based manipulation, (2) SAE-derived feature steering, and (3) probe-based direction steering. Overall, our results show that SAE- and probe-based techniques often outperform basic prompt-based methods for steering LLM agents, although this advantage narrows when compared with more advanced prompting strategies, such as few-shot and chain-of-thought prompting. Together, SAEs and probes constitute an effective pipeline for social scientists seeking to interpret and steer agents in social simulations: SAEs decompose agents' internal representations into human-readable features, after which probes can reliably shift agents' behaviors in specified directions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.