acceptodds
Under review as a conference paper at ICLR 2027

ReFrame: Learning Source-Conditioned Re-Embedding for Prompt Injection Defense

Abstract

Large language models (LLMs) follow trusted user instructions and leverage untrusted external data to complete tasks. However, malicious instructions embedded within such data can hijack model behavior through prompt injection. An effective defense must prevent such instructions from redirecting the model while preserving its ability to utilize information from external data. To address this challenge, we introduce ReFrame, a prompt-injection defense that separates user instructions from external data at the input representation level. First, source-conditioned orthogonal re-embedding rotates external-data embeddings while leaving user-instruction embeddings unchanged, preserving geometric relationships among external-data embeddings while reorienting them in the model's input space. Second, privilege calibration learns these rotations through preference and task-completion objectives, reinforcing instruction priority across input regions while maintaining task utility—all with pretrained model parameters frozen. Across four model backbones, ReFrame raises the average aggregate defense success rate from 43.14% to 88.13% and the average task utility score from 61.18 to 64.55 compared with unmodified models, with utility evaluated across attacked-task completion, document question answering, user-instruction execution, and code generation. These results demonstrate that learning input-space transformations over external data is an effective strategy for improving prompt-injection robustness without sacrificing task utility.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.