PRIVASIS ENTERPRISE: Learning Need-to-Know Text Sanitization from the Sharing Context
Abstract
To fix an API bug, an external coding assistant needs the stack trace, not the unreleased product described alongside it; an internal release review needs exactly that product. In need-to-know sanitization, a model rewrites records for a recipient and purpose, keeping what the purpose requires and withholding what the recipient should not see. Existing benchmarks spell out what to remove or keep, never testing whether a model can infer this from the request or reverse it for a new recipient. We build , a benchmark and training corpus of synthetic business records whose fact ledger tracks what can be inferred, so indirect leaks count. Paired recipients of the same records need opposite treatment of a fact, and Swap credits a model only when it gets both right. With category-level directives, the strongest general-purpose models look safer when judged by leakage alone, but most of their leak-free outputs also delete the necessary facts from the recipient. Compact sanitizers trained on our corpus meet every requirement more often than these models. Without directives, almost no output from any system meets every requirement, and Swap drops for every general-purpose model. Training on our corpus improves every metric but leaves this gap open, and its per-example objective never compares paired rewrites. Most missed swaps give both recipients the same answer: current sanitizers appear to judge what is sensitive, not who needs to know.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.