Stochasticity in Tokenization Improves Robustness
Abstract
The widespread adoption of large language models (LLMs) has increased concerns about their robustness. Vulnerabilities to perturbations of the input tokenization indicate that models trained with a deterministic canonical tokenization can be brittle to adversarial attacks. Recent studies suggest that stochastic tokenization can deliver internal representations that are less sensitive to perturbations. In this paper, we analyse how stochastic tokenizations affect robustness to adversarial attacks and random perturbations. We systematically study this over a range of learning regimes (pre-training, supervised fine-tuning, and in-context learning), datasets, and model architectures. We show that pre-training and fine-tuning with uniformly sampled stochastic tokenizations improve robustness to random and adversarial perturbations. Evaluating on uniformly sampled non-canonical tokenizations reduces the accuracy of a canonically trained Llama-1b model by 29.8%. We find that training with stochastic tokenization preserves accuracy without increasing inference cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.