The Refraction Gap: How Optical Context Compression Inadvertently Bypasses Safety Alignment
Abstract
To operate within finite context windows, interactive agents increasingly rely on optical context compression by rendering extensive textual inputs and interaction histories into compact visual pages. Existing evaluations of these transformations focus predominantly on capability preservation, examining whether models comprehend and act upon compressed content, while overlooking whether safety alignment is similarly maintained. We show that safety behaviors undergo a severe collapse across this representation boundary. Across five multimodal Readers and representative context compression pipelines, we reveal the Refraction Gap: harmful requests consistently refused in plain text are readily fulfilled when identical content is delivered in optical form. Targeted evaluations demonstrate that this vulnerability cannot be attributed to perceptual degradation: optical contexts preserve semantic fidelity and downstream task utility within 5.1 percentage points of the raw-text baseline, yet trigger unsafe compliance in an average of 47.5% of cases under payload rendering (reaching 73.1%) and 42.8% under memory compaction, compared to merely 4.5% in plain text. Finally, we show that this gap is mitigable prior to generation, where a lightweight 32M-parameter router identifies at-risk payloads before inference and routes them back to text. Taken together, our findings show that optical context compression is not merely an efficiency mechanism; it is a safety-relevant representation transformation that existing evaluation protocols fail to capture.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.