When Shopping Agents Are Misled: Understanding and Mitigating Purchase Safety Failures under Deceptive Merchants
Abstract
Shopping agents are increasingly asked to make purchase decisions, but existing e-commerce benchmarks primarily evaluate task completion in benign shopping environments. However, real-world shopping often involves deceptive merchants whose evidence and dialogue can lead agents to unsafe purchases. To bridge this gap, we introduce SafeShopBench, a purchase-safety benchmark for evaluating shopping agents in adversarial merchant environments. Experimental results demonstrate that SafeShopBench poses a substantial challenge even for frontier shopping agents. In our full-tool evaluation, unsafe-purchase rates across the twelve buyer agents range from 14.3% to 51.8%, and even the best-performing agent on this metric, Claude Opus 5, makes unsafe purchases in 14.3% of adversarial cases, highlighting the difficulty of purchase safety under deceptive merchants. To mitigate these purchase-safety failures, we evaluate two complementary approaches: a tool-grounded Detector-Corrector wrapper that revisits purchase decisions without retraining the buyer, and supervised fine-tuning followed by outcome-based reinforcement learning. For a smaller buyer agent, reinforcement learning reduces the unsafe-purchase rate from 24.9% after supervised fine-tuning to 8.3%, below that of Claude Opus 5. These results suggest that SafeShopBench supports both inference-time correction and the training of safer shopping agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.