acceptodds
Under review as a conference paper at ICLR 2027

Invisible Unicode Backdoors in Multilingual Seq2Seq Models: An Empirical Study across Language Resource Levels

Abstract

Multilingual pre-trained models enable NLP systems across hundreds of languages, yet their security vulnerabilities may differ substantially across resource levels. We present an empirical study of invisible Unicode backdoor attacks on multilingual seq2seq models spanning high-, mid-, and low-resource languages. Across MT0, mT5, UMT5, and ByT5, we observe a consistent resource-level vulnerability gradient: low-resource languages achieve higher attack success rates (ASR) at comparable poison rates and require substantially lower minimum effective poison rates (MEPR) than high-resource languages. We further examine cross-lingual transferability, tokenization fragmentation, model scale, and experimental setting as factors associated with this asymmetry. Cross-lingual transferability is strongly correlated with monolingual backdoor susceptibility (r = 0.931, p < 0.01), while tokenization fragmentation amplifies the attack effect in T5-family models. Experiments with the byte-level ByT5 model show that the vulnerability gradient persists even without subword tokenization fragmentation, suggesting that fragmentation is an amplifying factor rather than the sole explanation. We also evaluate a lightweight input sanitization strategy based on Unicode normalization and invisible-character filtering, which substantially reduces attack success while causing only small changes in clean task performance in the tested settings. These findings highlight the need for multilingual-aware security evaluation and defense strategies, particularly for lower-resource languages.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.