What the Model Sees: A Safety Evaluation of LLM-based Email Assistants in Real Mail Clients
Abstract
Large language models are increasingly used for daily tasks, including email processing and action. Existing evaluations often test phishing detection from fixed inputs, while deployed AI assistants may receive emails only after mail clients render and transform them. We study 21 mail clients, 16 LLMs, four input forms, and six AI browser extensions. We define a run as safe when the assistant recognizes the risk, warns the user, and refuses to act on the message. We separate message accessibility, risk judgment, and user protection, which are often conflated in prior evaluations. Safe rates across 16 LLMs range from 28.6% to 100.0%, while message accessibility across 21 clients ranges from 16.7% to 97.8%. Input representation also matters: the same attack message reaches a safe verdict in 95.0% of screenshot runs versus 54.0% with HTML uploads. Client warnings increase the safe rate from 57.5% to 86.1%, while a lightweight risk instruction increases it from 68.1% to 81.4%. Across six AI browser extensions, only 24.4% of delegated email tasks result in a warning or refusal. These results show that safe email delegation depends on message exposure, input representation, and contextual safety signals. We recommend that mail services and clients provide AI assistants with both risk assessments and readable message content.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.