acceptodds
Under review as a conference paper at ICLR 2027

The Refraction Gap: How Optical Context Compression Inadvertently Bypasses Safety Alignment

Abstract

To operate within finite context windows, interactive agents increasingly rely on optical context compression by rendering extensive textual inputs and interaction histories into compact visual pages. Existing evaluations of these transformations focus predominantly on capability preservation, examining whether models comprehend and act upon compressed content, while overlooking whether safety alignment is similarly maintained. We show that safety behaviors undergo a severe collapse across this representation boundary. Across five multimodal Readers and representative context compression pipelines, we reveal the Refraction Gap: harmful requests consistently refused in plain text are readily fulfilled when identical content is delivered in optical form. Targeted evaluations demonstrate that this vulnerability cannot be attributed to perceptual degradation: optical contexts preserve semantic fidelity and downstream task utility within 5.1 percentage points of the raw-text baseline, yet trigger unsafe compliance in an average of 47.5% of cases under payload rendering (reaching 73.1%) and 42.8% under memory compaction, compared to merely 4.5% in plain text. Finally, we show that this gap is mitigable prior to generation, where a lightweight 32M-parameter router identifies at-risk payloads before inference and routes them back to text. Taken together, our findings show that optical context compression is not merely an efficiency mechanism; it is a safety-relevant representation transformation that existing evaluation protocols fail to capture.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.