acceptodds
Under review as a conference paper at ICLR 2027

Under the Influence: Evaluating and Improving LLM Vigilance in Agentic Tasks

Abstract

Large Language Models (LLMs) are becoming increasingly autonomous assistants, capable of accomplishing series of complex tasks by interacting with open world domains such as software databases and the internet. This presents great opportunity but also great risk: open domains routinely present opportunities for autonomous agents to be influenced by others, e.g. deciding whether to accept pull requests or whether to use advice on a message board in order to accomplish a task. Critically, agents must be capable of deciphering when to incorporate benevolent advice into their reasoning, and when to reject malicious advice. This involves the social capacity of vigilance: the ability to determine which information to use, and which to discard, for effective communication. While well-studied in humans, the extent to which LLM agents exhibit vigilance is relatively unknown. Here, we examine LLM agent vigilance in three realistic tasks: navigating dangerous websites, avoiding e-commerce scams, and patching an unsafe code repository. We find that all frontier models are highly susceptible to malicious influence across all environments. We also find that model performance itself is surprisingly dissociable from vigilance. In light of these results, we propose a vigilance harness which uses Bayesian reasoning over a model's estimates of the value of an advisor's advice to decide whether to act on that advice. The vigilance harness helps recover baseline performance while retaining the benefit of positive influence on agents, paving the way for safer agentic AI systems.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.