DepthGate: Depth-Aware Gated Steering for Instruction-Data Separation in LLMs
Abstract
Large language models often process trusted instructions that specify the task and untrusted data that supply task-relevant content within the same context. Prompt injection arises when adversarial instructions embedded in untrusted data override the trusted task. Instruction-data separation makes this role boundary explicit, but existing approaches enforce it through system-level controls or input-level role encoding, leaving its treatment during depth-wise model computation largely unspecified. We introduce DepthGate, a depth-aware framework that extends instruction-data separation from input formatting to hidden-state computation, carrying the trusted role boundary into depth-wise model computation. DepthGate achieves this through three complementary mechanisms. Role-Conditioned Steering introduces continuous role-aware gates to control the redirection of data representations, while keeping instruction representations unchanged and preserving the intrinsic geometry of data. Depth-Aware Paths organize these gates across model layers, providing model-specific control over where and how strongly the separation boundary is enforced during forward computation. Schedule Discovery further identifies steering schedules under schedule-matched adaptation and selects a model-specific schedule using held-out data before final post-training. Experiments across five Llama and Qwen models show that DepthGate consistently improves instruction-data separation, with its most consistent reductions in marked-data attack success. Different schedule variants further provide flexible robustness-utility trade-offs. On Qwen2.5 7B, the high-separation configuration raises the separation score from 51.0% to 92.1% and reduces StruQ-ID attack success from 72.4% to 11.9%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.