acceptodds
Under review as a conference paper at ICLR 2027

Interpreting and Steering LLM Agents for Social Simulations

Abstract

Generative agents built on large language models (LLMs) have proven powerful for simulating human behavior, making them valuable additions to the social scientific toolkit. However, the black-box nature of LLMs limits their value for simulation by constraining both: (i) interpretability: uncovering mechanisms driving observed behavior; and (ii) steering: muting or amplifying mechanisms of action along theoretically meaningful dimensions. Here, we demonstrate how the black box could be opened up to further enrich LLM-based simulations. We illustrate our approach by interpreting and steering two primitive components of human behavior, namely preferences (risk attitudes, fairness preferences) and capabilities (divergent creativity, product innovation), using four classic economic and creative tasks implemented as natural-language interactions. We demonstrate the utility of sparse autoencoder (SAE)-based techniques for interpreting mechanisms and compare three types of steering methods: (1) prompt-based manipulation, (2) SAE-derived feature steering, and (3) probe-based direction steering. Overall, our results show that SAE- and probe-based techniques often outperform basic prompt-based methods for steering LLM agents, although this advantage narrows when compared with more advanced prompting strategies, such as few-shot and chain-of-thought prompting. Together, SAEs and probes constitute an effective pipeline for social scientists seeking to interpret and steer agents in social simulations: SAEs decompose agents' internal representations into human-readable features, after which probes can reliably shift agents' behaviors in specified directions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.