AI agents are susceptible to social influence
Abstract
AI agents increasingly act in environments that expose them to text from parties other than their principal (e.g., shared documents, online forums or code-review threads). Here we ask which properties of such text, explicitly attributed to a third-party, makes it more likely to influence agent behaviour. To this end, AI agents solving GitHub issues (adapted from SWE-bench items) were provided with a comment thread showing suggestions to use a shortcut (hack) to complete the task. Across thirteen AI models, we find that influence on agent behaviour is 1) stronger when the comment is attributed to a human rather than an automated assistant; 2) stronger when the comment mentions time pressure; and 3) unaffected by the level of agreement of other AI agents. These results show that agents can take actions their principal did not intend in response to ordinary third-party comments, and that how such comments are presented matters. This should be accounted for when deploying AI agents in online environments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.