DARE: Efficient Multilingual PII Detection
Abstract
Detecting personally identifiable information (PII) is a fundamental component of privacy-preserving natural language processing, requiring models that reliably identify sensitive information across languages and heterogeneous data distributions. However, current models either involve billions of parameters and achieve-state-of-the-art performance, or involve millions of parameters and fall short in performance. In this work, we introduce DARE, an efficient bidirectional encoder for multilingual PII detection with 300M parameters, trained on a heterogeneous mixture of publicly available datasets spanning multiple languages and annotation schemes. DARE achieves strong detection performance across well-established multilingual privacy benchmarks and outperforms state-of-the-art open-weight PII detection baselines with up to more total parameters under a common evaluation protocol, while incurring up to lower inference overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.