UAL Semantic Guard (USG): Efficiently detecting UAL prompt leakage attacks on LLMs
Abstract
Several black-box mechanisms protect large language models (LLMs) against prompt data leakage, an attack that forces a model to leak data derived from users’ private prompts. However, we found that they are extremely vulnerable to user attribute leakage (UAL) inference. Unlike other prompt data leakage attacks, UAL inference specifically targets user data that is not explicitly provided but can be inferred. Our evaluation shows that, against such attacks, the best overall protection mechanism, Meta PromptGuard, drops to only a 5% detection rate. To tackle this problem, we introduce UAL semantic guard (USG), a black-box prompt data leakage detector that also efficiently detects UAL inference. USG relies on a key idea: augmenting a small model’s context and providing attack and benign examples improve its ability to detect prompt data leakage attacks. We leverage this idea and augment a local model to act as an LLM-as-a-Judge that blocks prompt processing if it can lead to prompt data leakage. We evaluate our prototype against 3600 adversarial queries spanning nine UAL-Inference variants and 3200 benign queries covering several user tasks. We also evaluate against several categories of prompt data leakage. We run each evaluation on four open-weight LLMs: Llama 2, Llama 3.1, Mistral-Nemo, and Qwen. Our results show that (1) USG achieves 100% detection across all UAL attack variants, (2) while maintaining a 1.19% false-positive rate on benign queries, (3) outperforming the best state-of-the-art approaches, which only reach 55.2%, (4) with no major impact on prompt pro- cessing time and energy usage. Additionally, it maintains a 100% detection rate against other prompt data leakage attacks, demonstrating overall robustness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.