acceptodds
Under review as a conference paper at ICLR 2027

Preserving Semantic Integrity: Who May Covertly Redefine an LLM's Task?

Abstract

Traditional prompt injection attacks add, replace, or override instructions that compete with what a language model is authorized to do. This paper studies a different problem, which we call semantic reanchoring. The authorized instruction remains literally unchanged, while untrusted examples or records change what parts of it are taken to mean. A pattern such as OFF -> ON can cause "turn OFF the stove when no one is home" to be interpreted and executed as turning it ON. Likewise, examples can reanchor "-" from subtraction to addition, turning a 100-10 discount into 110. Across more than eighty-five thousand generations and verification calls spanning eight LLMs, we show that preserving literal or instruction integrity does not guarantee semantic integrity: a model can coherently execute an unauthorized interpretation of an unchanged instruction. We make authority explicit by testing both cases in which contextual reinterpretation is forbidden and cases in which the owner authorizes it, separating unauthorized reinterpretation from legitimate adaptation and rigidity. Reasoning mode and owner wording can move behavior in opposite directions across configurations, and an interpretation can later be restored toward its prior meaning, complicating retrospective detection. Crucially, semantic verification is feasible. An LLM auditor given the owner's instruction, the proposed action, and the complete owner-authorized semantic reference detects nearly all observed violations across both audited task families with low false rejection. Surprisingly, hiding contextual material from the auditor can make verification substantially worse when the owner delegated meaning to that material. The defense boundary is therefore not context versus isolation, but authorized reference versus unauthorized influence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.