acceptodds
Under review as a conference paper at ICLR 2027

PRIVASIS ENTERPRISE: Learning Need-to-Know Text Sanitization from the Sharing Context

Abstract

To fix an API bug, an external coding assistant needs the stack trace, not the unreleased product described alongside it; an internal release review needs exactly that product. In need-to-know sanitization, a model rewrites records for a recipient and purpose, keeping what the purpose requires and withholding what the recipient should not see. Existing benchmarks spell out what to remove or keep, never testing whether a model can infer this from the request or reverse it for a new recipient. We build , a benchmark and training corpus of synthetic business records whose fact ledger tracks what can be inferred, so indirect leaks count. Paired recipients of the same records need opposite treatment of a fact, and Swap credits a model only when it gets both right. With category-level directives, the strongest general-purpose models look safer when judged by leakage alone, but most of their leak-free outputs also delete the necessary facts from the recipient. Compact sanitizers trained on our corpus meet every requirement more often than these models. Without directives, almost no output from any system meets every requirement, and Swap drops for every general-purpose model. Training on our corpus improves every metric but leaves this gap open, and its per-example objective never compares paired rewrites. Most missed swaps give both recipients the same answer: current sanitizers appear to judge what is sensitive, not who needs to know.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.