PII-RADAR: RApid Detection And Replacement for PII Scrubbing
Abstract
Large language models have been shown to leak their training data, which becomes a privacy risk when this data contains personal information. In practice, vulnerable data is commonly removed using scrubbers, *i.e.*, tools that detect personally identifiable information (PII) in a document and remove it before training. However, scrubbers are imperfect and miss some PII in the training data. In this work, we first show that, while scrubbing leaves less information exposed, the PII it misses becomes significantly more vulnerable. The leakage we observe increases with the detection rate, so scrubbers with stronger detectors increase the risk for missed PII. We identify PII removal as the primary cause of this leakage, because it leaves only a small amount of real PII in the dataset, which the model then memorizes. While synthetic rewriting of the detected PII should, in theory, resolve this issue, current rewriting tools generate highly repetitive or malformed text, increasing the risk similarly to removal. To address this issue, we propose PII-RADAR, a novel scrubber that replaces each detected PII with a realistic one of the same type, drawn from a pool of validated synthetic identities and formatted to match the original document. Contrary to other rewriting scrubbers, PII-RADAR's replacements are diverse, well-formed, and non-repetitive, lowering the privacy risk of any PII it misses **up to 3×** compared to baselines. Moreover, PII-RADAR outperforms all evaluated scrubbers on five of seven benchmarks, establishing a new state of the art in PII detection, while being up to 10× faster than other scrubbers that replace rather than redact.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.