acceptodds
Under review as a conference paper at ICLR 2027

Mitigating Privacy Leakage in Text Corpora via Adaptive Masking and Contextual Reconstruction

Abstract

Large-scale, high-quality data is central to advances in artificial intelligence, yet raw corpora are riddled with private information, exposing their direct use to serious security and privacy risks. Existing methods struggle to reconcile semantic preservation with private content removal, thereby limiting both privacy protection and downstream data utility. In this paper, we introduce PrivMask, a method that mitigates privacy leakage in text corpora via adaptive masking and contextual reconstruction, while preserving task-relevant semantics and downstream utility. Specifically, we first propose a category-aware sensitive proportion masking mechanism that adaptively masks sensitive content by modeling category-level sensitivity and contextual risk. Instead of using coarse placeholders, it abstracts privacy entities into semantically meaningful labels, enabling fine-grained removal of sensitive information while preserving semantic structure and preventing over-sanitization. We then present a context-aware reconstruction model that generates fluent and semantically consistent substitutes from masked text. To simultaneously achieve grammatical coherence and fine-grained semantic alignment, we optimize the model with a joint loss that integrates reconstruction and contrastive objectives. This joint training ensures that the reconstructed substitutes retain task-relevant semantics without leaking the original private data, achieving a robust privacy-utility trade-off. Extensive experiments across diverse domains show that PrivMask outperforms existing methods, achieving high accuracy in privacy entity recognition while preserving semantic coherence. Notably, models trained on PrivMask-anonymized data retain competitive downstream performance, confirming that PrivMask effectively balances privacy protection and data utility.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.