acceptodds
Under review as a conference paper at ICLR 2027

PRISM: Transparent Strategy Search for Privacy Risks in LLM Agents

Abstract

Discovering privacy attacks on LLM agents by search is expensive: each candidate attacker instruction must be evaluated in a simulated multi-agent conversation, and the search framework of Zhang & Yang (2026) decides what to try next inside one reflect-and-rewrite LLM call, with no explicit representation of tactics and no surrogate-based rule for choosing among proposals before they are simulated. PRISM is a drop-in replacement for that proposal step. It represents each instruction as a sparse code over named tactic primitives learned from its text, fits a closed-form Bayesian surrogate of leak velocity (the benchmark’s metric; higher means a stronger attack), and uses the surrogate to choose which previously evaluated instructions the parallel search threads revise next; the rewrite itself remains the baseline’s own LLM call, and the simulator, judge, metric, data splits and per-step simulation and optimizer-call budgets are unchanged. On the benchmark’s 100 held-out scenarios, run entirely on open-weight models, the attack instructions transferred from PRISM’s search reach a mean leak velocity of 0.277 against 0.188 for the baseline’s (paired difference +0.089, 95% bootstrap interval [+0.025, +0.153] over scenarios). We also measure a property of the benchmark’s selection rule, which breaks ties among equally scored candidates uniformly at random: from a defended starting point in our environment all 30 candidates had zero observed leakage in 87 of 100 steps, so the observed scores did not distinguish candidates within those steps. A game-theoretic defense construction is derived and evaluated offline on released trajectories; the comparative defense searches stopped before it could activate, so no live comparison of defense proposers is reported.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.