acceptodds
Under review as a conference paper at ICLR 2027

Your Agentic LLMs Can Latently Detect Indirect Prompt-Injection Attacks

Abstract

Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, \eg, malicious side-tasks hidden in external tool results. While many efforts have sought to address this threat, little is known about the internals of these LLMs when they are exposed to IPI attacks (named IPI exposure for simplicity). In this paper, we study IPI exposure from three perspectives. (1) Probing: Across eight models, including 753B GLM-5.2 and 2.8T Kimi-K3, simple linear probes trained on LLMs' pre-generation hidden states can predict their IPI exposure. These probes achieve 0.90+ AUROC on unseen attacks, agent instructions, and task suites; notably, they can remain robustly predictive under adaptive attacks and in cross-lingual settings. (2) Defense: Our CoT-monitoring analysis diagnoses the knowledge–action gap: these LLMs often fail to reliably bind these signals to safe agentic actions. We therefore introduce a probe-gated reasoning-based defense to bridge this gap at test time. On difficult AgentDojo settings, the simple defense can substantially reduce attack success rate, \eg, from 34.6% to 0% on Qwen3.5-27B, and better preserves clean-task utility than the baselines. (3) Explanation: We introduce an analysis framework that identifies natural-language explanations strongly correlated with probe-captured signals. The resulting profiles differ across probes trained on different model-layer combinations: latent signals can align with either direct IPI-exposure sensing or indirect operational cues. Code is available: https://anonymous.4open.science/r/IPI-exposure.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.