acceptodds
Under review as a conference paper at ICLR 2027

PrivacyArena: Evaluating AI Agents on Privacy Attacks Against ML Models

Abstract

AI agents are increasingly used for security auditing of machine learning models, but no existing benchmark measures whether agents can autonomously develop and execute privacy attacks. This capability is dual-use: it can enable automated red-teaming of deployed models, but may also lower the barrier to misuse as agents become more capable, and should therefore be carefully assessed. We present PrivacyArena, a benchmark for assessing agents' autonomous privacy-attack capabilities across 151 tasks spanning seven types of privacy-relevant attacks: membership inference, canary extraction, attribute inference, property inference, dataset inference, data reconstruction and model stealing across vision, tabular, and language models. Each task provides instructions, target model(s), and a sandboxed GPU environment in which agents must perform an attack from scratch. We also re-implement state-of-the-art attacks for every task under the same compute budget the agent receives, giving a calibrated reference for what a competent researcher could achieve. We evaluate three frontier agents, Claude Code with Claude Opus 4.7, Codex with GPT 5.5, and Gemini CLI with Gemini 3.1 Pro, on all tasks, and ablate model and harness choice on a task subset across 12 additional models and 6 harnesses. We find that agents can perform most of the attacks we investigate. They do not consult any resources about the attacks, instead relying on prior knowledge, and often stop early without offline validation. Catastrophic failures still occur, especially when agents are confused about the data. We publish the benchmark at https://anonymous.4open.science/r/privacy-iclr27-86A3 to support research on automated auditing and on tracking agent attack capabilities.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.