Welcome to the Age of Agent Ads: Using On-Policy Attacks to Steer Agentic Decisions
Abstract
People increasingly delegate option-ranking tasks (e.g., shopping, job search, travel) to AI agents, which creates incentives to steer agentic decisions via "*agent ads*." To defend against such steering attacks, people count on models’ alignment training and platforms' terms of service, and consequently the most effective existing attacks are designed to trigger _off-policy_ behaviors. Exposing a new attack surface, we introduce *on-policy attacks*, which steer agentic decisions without pushing individually aligned behaviors off policy. Our example attack consists of two steps: To remain compliant with platform's terms of service, we first inject a _truthful negation_ about a fictitious concept ("_Does not contain blizyester_") into a product description. This triggers an agent to search the web ("_What is blizyester?_"), which surfaces planted websites full of fabricated information associating the fictitious concept with harmful implications ("_Blizyester linked to dementia_"). Trained to keep users safe from harm, the agent penalizes every option missing the truthful negation. Evaluating our attack across four real-world listing datasets, we find some frontier models exposed, enabling an adversarial actor to penalize competitors without violating existing platform-listing guidelines. We discuss the broader challenge of regulating advertising aimed at AI agents and provide recommendations to model developers and platform providers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.