acceptodds
Under review as a conference paper at ICLR 2027

Welcome to the Age of Agent Ads: Using On-Policy Attacks to Steer Agentic Decisions

Abstract

People increasingly delegate option-ranking tasks (e.g., shopping, job search, travel) to AI agents, which creates incentives to steer agentic decisions via "*agent ads*." To defend against such steering attacks, people count on models’ alignment training and platforms' terms of service, and consequently the most effective existing attacks are designed to trigger _off-policy_ behaviors. Exposing a new attack surface, we introduce *on-policy attacks*, which steer agentic decisions without pushing individually aligned behaviors off policy. Our example attack consists of two steps: To remain compliant with platform's terms of service, we first inject a _truthful negation_ about a fictitious concept ("_Does not contain blizyester_") into a product description. This triggers an agent to search the web ("_What is blizyester?_"), which surfaces planted websites full of fabricated information associating the fictitious concept with harmful implications ("_Blizyester linked to dementia_"). Trained to keep users safe from harm, the agent penalizes every option missing the truthful negation. Evaluating our attack across four real-world listing datasets, we find some frontier models exposed, enabling an adversarial actor to penalize competitors without violating existing platform-listing guidelines. We discuss the broader challenge of regulating advertising aimed at AI agents and provide recommendations to model developers and platform providers.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.