Differentially Private Byte Pair Encoding Tokenizers for Large Language Models
Abstract
Tokenizers are a standard part of large language model (LLM) releases, providing practical tools for splitting text and counting tokens. However, vocabulary and ordered merge list released with common byte pair encoding (BPE) tokenizers for LLMs reflect recurring patterns in the data corpus used during their training. We ask whether access to such a tokenizer can reveal that a particular website was included in its training data. Under three access models, exposing (1) token counts only, (2) token identities, and (3) the full BPE tokenizer with both the vocabulary and the ordered merge list, we design website-level membership inference attacks calibrated using shadow tokenizers trained with and without the target website. Under the most restrictive token-count access, our attack reaches 60.9 TPR at 1% FPR, while the best prior attack under the same access stays at 1.3% and therefore does no better than guessing. Motivated by this leakage, we also evaluate defenses. We find that simple low-count filtering reduces the attack success under token-count access but does not reliably protect against the stronger token-identity and merge-list attacks. Moreover, prior defenses provide no provable guarantees on privacy leakage. We therefore propose DP-BPE, the first differentially private BPE training algorithm, which gives the released vocabulary and merge list a provable website-level privacy guarantee. At a 4% relative utility cost, it reduces every attack we evaluate to close to random guessing, while prior defenses leak substantially more. We also open-source our DP-BPE as a Rust implementation for the community. Together, our findings highlight the need for privacy-aware tokenizer training and show that a website-level privacy guarantee is possible at a small utility cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.