acceptodds
Under review as a conference paper at ICLR 2027

Adversarial Accuracy Misses What Linear Probes Still Infer from Anonymized Text

Abstract

Large language models infer personal attributes from online comments with near-human accuracy, and text anonymizers rewrite the comments to prevent such inference. Anonymizers are evaluated by adversarial accuracy, the share of attributes that an LLM attacker still infers correctly. Most of them also use such an attacker while rewriting, so a lower accuracy may show that the attacker's answer has changed rather than that the attribute has left the text. We therefore examine the text with an adversary outside the rewriting loop, a linear probe trained on other authors' labeled raw text. On SynthPAI, six published anonymizers remove up to all of the attacker's advantage over a majority-class guess, but at most 15% of the probe's advantage over chance. The gap remains for retrained probes and a fine-tuned classifier, and a frozen probe finds it on real Reddit comments. Controlled edits on SynthPAI suggest why: much of the remaining signal lies in register, how a comment is written rather than what it states. To narrow the gap, Probe-Guided Register Selection (PGRS) rewrites each comment in several registers and publishes the rewrite the defender's own probe is least certain about. PGRS lowers adversarial accuracy more than most published anonymizers, at a moderate utility cost. PGRS lowers frozen-probe AUC from 0.76 to 0.51 on three probe models, two of which take no part in selection. Because it selects against a frozen probe, we also report a probe retrained on its output, which reaches 0.56-0.60 against 0.73-0.77 for the published anonymizers. We therefore recommend reporting probe AUC, frozen and retrained, alongside adversarial accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.