MuPPET: A Benchmark for Contextual Privacy of LLM Assistants in Multi-Party Conversations
Abstract
LLM agents are increasingly deployed in multi-party environments, such as group chats, handling sensitive personal data on behalf of individual users. When an agent communicates, the message reaches every group member at once. The model must therefore judge whether disclosing private information is appropriate for every recipient simultaneously, a risk that is structurally harder to control than in one-to-one conversations. Yet existing contextual privacy benchmarks focus on one-to-one settings, so multi-party privacy risks remain largely unmeasured. We introduce MuPPET (Multi-Party Privacy Exposure Testing), a benchmark for contextual privacy in multi-party conversations. Our experiments show that multi-party conversations increase leakage most for the models that appear safest in one-to-one conversations, by up to a factor of 1.9. Small open-weight models leak the most in both settings, a concern since they are often preferred for local deployment with sensitive data. Existing contextual privacy defences offer only partial protection, degrade utility, and do not resolve the underlying party-tracking problem.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.