ĪCÁRŪS: Flying Over Safety Alignment to Jailbreak Large Language Models via Ancient Grammar
Abstract
Large language models are trained to refuse harmful requests, yet jailbreak attacks still bypass this safety alignment. We show that expressing a harmful request in an ancient language reliably bypasses alignment, and that the effect is carried by grammar rather than vocabulary, concentrated in its orthographic rules. The most effective grammar rules fragment harmful words into many tokens and shift the input embedding away from English, while barely changing the visual form of the prompt. Building on this, we propose Iterative Combinatorial Ancient Rule Search ĪCÁRŪS, a black-box attack that uses beam stack search over a pool of ancient grammar rules to find an effective combination for each request. Across five open- and closed-source models, ĪCÁRŪS outperforms strong baselines on StrongREJECT and attack success rate, and its plain-English output is directly usable. On a widely-used model, an ancient-language translation alone raises the attack success rate from 15% to 48%, and ĪCÁRŪS raises it as high as 85%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.