acceptodds
Under review as a conference paper at ICLR 2027

MSO-Bench: Benchmarking and Detecting Malicious Semantic Overreach in Documents

Abstract

Embedded content in untrusted documents can steer LLM decisions and downstream actions. A natural defense is to use a lightweight LLM to pre-screen documents for unsafe content before they reach downstream systems. However, existing detectors largely rely on surface cues, missing attacks expressed as ordinary document content while flagging benign documents that contain prompt-like language but pose no actual threat. We therefore ask whether a lightweight detector can distinguish malicious from benign content by considering what function a segment serves in its document context and whether the actions it induces are indeed malicious. We introduce MSO-Bench, a benchmark spanning 20 real-world document types, that varies surface cues independently of the document's function and induced action. Evaluation results show that existing detectors perform poorly on MSO-Bench, revealing substantial gaps in current document screening. We subsequently train MSO-Guard, a 4B detector designed to reason about contextual function and induced actions to make document-level safety decisions. MSO-Guard achieves a 94.2% F1 score, substantially outperforming existing dedicated detectors and matching GPT-5 at a fraction of the deployment cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.