Do Not Assume Users Are Privacy-Aware: Evaluating Proactive Privacy Risk Warning in Large Language Models
Abstract
The widespread deployment of large language models (LLMs) has made safety a critical research priority. Existing studies predominantly focus on adversarial threats, such as harmful generation, and model-side privacy leakage. However, this assumes risk originates from malicious attackers, overlooking a common real-world scenario: benign users may unintentionally include sensitive information even when making legitimate requests. Such inclusion creates privacy risks that LLMs should proactively recognize, even without explicit privacy instructions. This calls for models to go beyond task completion by alerting users to these risks and preventing unintended data propagation. To evaluate this issue, we introduce PrivAwareBench, a benchmark designed to evaluate LLMs’ awareness of privacy risks arising from benign users’ unintentional disclosure. It assesses how this awareness is reflected in two observable behaviors: risk warning and content suppression, across diverse exposure conditions. The benchmark data are constructed with human involvement and manual quality control, and the evaluation judge is selected through a human–LLM agreement study. Experiments on 18 representative LLMs reveal that proactive privacy-aware behavior remains limited, category-sensitive, and brittle. These results expose a blind spot in protecting benign users from unintended disclosure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.