acceptodds
Under review as a conference paper at ICLR 2027

AttackRecBench: Benchmarking Item Promotion Attacks on Agentic Recommender Systems

Abstract

Large language models are transforming recommender systems from one-shot rankers into agents that retrieve external information and maintain state across multi-step decisions. Existing recommendation-attack benchmarks do not systematically evaluate the new attack surfaces introduced by these agentic capabilities. In particular, attacker-controlled item titles, descriptions, and public reviews can manipulate evidence retrieval, reasoning, memory, and final rankings. We introduce AttackRecBench, a benchmark for item-promotion attacks on agentic recommender systems. It comprises 1,500 candidate-set ranking tasks from Amazon, Goodreads, and Yelp across five recommendation scenarios. The benchmark includes eleven attacks spanning content rewriting, indirect prompt injection, decision-bias manipulation, and working-memory poisoning, evaluated across five recommender-agent architectures. Attackers can modify only seller- or user-controlled content, without access to model parameters, user histories, system instructions, or candidate sets. Paired clean and adversarial evaluation measures target-item promotion and recommendation utility. In our primary evaluation setting, prompt injection is the strongest promotion mechanism, increasing the top-three exposure rate by 17.7–40.6 percentage points across the three domains. Attack effectiveness is nevertheless conditional: the strongest individual attack changes across recommendation scenarios and agent architectures, while responses to the memory-oriented attack condition vary markedly across architectures. We further evaluate three lightweight inference-time defenses. All three reduce the top-three exposure rate of the strongest contextual prompt-injection attack by 25.6–29.7 percentage points, but their security and utility effects vary across attack families; no defense dominates overall.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.