SEA: Toward Safety-Aware AI Agents in Realistic Social Engineering Attack Scenarios
Abstract
Personal AI agents are increasingly deployed on user devices to execute actions involving real money and sensitive data, but their security under social engineering attacks has not been systematically evaluated. We propose SEA (Social Engineering Attack Benchmark), the first interactive environment for evaluating the intrinsic security capability of AI agents in social engineering attack scenarios, covering 8 daily-life domains, 30+ simulated applications, a 4-level difficulty gradient, and a 5-dimensional evaluation framework, with 320 evaluation tasks designed in consultation with anti-fraud specialists and spanning two culturally distinct settings in Chinese and English. Evaluation on 11 frontier open-source and closed-source models, with GPT-6 evaluated only in Chinese, reveals that all models are breached to varying degrees (Chinese attack success rates ranging from 16.9% to 84.3%), with maximum single-conversation losses exceeding ¥ 40,000, proactive verification rates universally below 32%. We further propose MetaSEA, a co-evolutionary framework for attack tasks and defenders. With MetaSEA training, Qwen3-8B achieves defense performance comparable to closed-source models while effectively reducing ungrounded action claims.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.