acceptodds
Under review as a conference paper at ICLR 2027

Hidden in Plain Sight: Advancing Harmful Content Understanding via Perception-Aware Adversarial Self-Play

Abstract

Despite significant advances in language understanding, Large Language Models (LLMs) can still miss meaning hidden in plain sight, failing to identify harmful intent in expressions readily understood by humans. Malicious users exploit this gap by strategically disguising harmful content to evade automated moderation, posing a serious threat to online platform safety. The vulnerability reflects a fundamental cognitive asymmetry between human readers, who integrate contextual, phonetic, and visual cues to decipher the intent behind deliberate obfuscation, and text-based LLMs, which rely primarily on token sequences and have limited exposure to evolving evasions. To bridge this asymmetry, we propose PASP (Perception-Aware Adversarial Self-Play), which combines multi-perceptual views with adversarial self-play to improve understanding of harmful intent as evasions evolve. Specifically, PASP constructs textual, phonetic, and visual views of each input, preserving its original context while exposing pronunciation and glyph cues obscured in token space. During adversarial self-play, an attacker uses defender feedback to generate intent-preserving evasions, while the defender is trained using an intent-invariance reward to understand harmful intent across changing surface forms. Extensive experiments show that PASP reduces adaptive attack success rates by up to 73.71% and 74.49% on ChineseHarm-Bench and EvoHarmBench, respectively, preserves model capabilities with minimal degradation, and effectively interprets real-world evasive expressions in e-commerce.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.