Telltale Circuits? A Controlled Comparison of Text, Activation, and Attribution-Graph Monitors for Covert LLM-Agent Sabotage
Abstract
A backdoored LLM agent can hide sabotage in one ordinary-looking tool call, with its trajectory remaining benign until the critical argument is emitted. We ask whether monitoring the agent's computation through attribution graphs yields a better or harder-to-evade detector than monitoring its text or activations. We compare three monitor families under a common protocol: text monitors (TF-IDF and a zero-shot LLM judge), a linear activation probe, and a structural monitor over attribution-graph fingerprints at the decision token. On trigger-gated LoRA model organisms in Gemma-3-4B, with key findings replicated on Qwen3-4B, attribution structure carries sabotage-correlated signal that a zero-shot text judge does not detect, and for exfiltration this signal survives a control that fixes the attribution target. At this scale, however, we find no detection advantage over the linear probe wherever our confound controls permit a clean comparison. Fixing the attribution target reduces pooled in-distribution structural AUROC from 0.826 to 0.420, and an apparently perfect repeated-mechanism generalization result is also captured by cheap residual-stream summaries, including a 50-component PCA. A perturbation that collapses the probe (0.780 to 0.228) leaves the structural monitor nearly unchanged when transferred (0.844 to 0.847) but also leaves an independently fit second probe at 1.000, so this non-transfer does not establish structural robustness. An attack aimed at the structural score, built from a one-hop surrogate and verified on the full pipeline, reduces AUROC to 0.639 on Gemma-3-4B and 0.559 on Qwen3-4B while sabotage behavior is preserved on retained examples. Because this attacker writes directly to the residual stream, the attack provides an upper bound on evasion rather than a realistic input-space adversary. We contribute a matched monitoring protocol and a map of the confounds and failure modes that can make structural monitors appear stronger than they are. Structural monitoring also costs substantially more compute than the baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.