acceptodds
Under review as a conference paper at ICLR 2027

Word Substitution Bypasses Safety Monitors Despite Explicit Mappings

Abstract

Word substitution can bypass language-model safeguards by replacing sensitive terms with harmless-looking words, even in some of the recent advanced models. External safety monitors add another layer of defense, yet it remains unclear whether they can recognize harmful content even when told what the substituted words mean. To further investigate this issue, we conducted a comprehensive and systematic study of this phenomenon. Our empirical study first reveals that, across seven generation services, substitution-based attacks persist against some frontier models; seven monitors show 13.3–76.7 percentage-point alarm reductions on paired held-out inputs despite explicit meanings; and restoring questions increases alarms at least as much as restoring answers. Also, in an open-weight monitor, swapping codeword meanings while fixing vocabulary, questions, and answers yields distinguishable hidden states but unchanged safety decisions: mappings are linearly separable on unseen questions, while scores shift but remain below threshold. Explicit harmful concepts instead trigger some alarms. On generated answers, templates, output prefixes, and general instructions receive 92% of direct decision-position attention, with similar aggregate allocations across mappings. Demonstrations improve detection on unseen questions, and lightweight fine-tuning improves risk ranking under matched inference settings. Identical warnings yield different harmful-delivery outcomes depending on whether triggering text is withheld. Our findings show that explicit word meanings and distinguishable internal states are insufficient for reliable safety monitoring, informing how external monitors should be evaluated and improved.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.