acceptodds
Under review as a conference paper at ICLR 2027

CompassGuard: Context-Aware Activation-Based Defense for Multi-Agent Systems

Abstract

LLM-based multi-agent systems enable collaboration through inter-agent communication, but these interactions can also propagate malicious instructions from compromised agents. Pruning-based defenses can disrupt collaboration, while prototype-based activation defenses use a shared benign reference that does not explicitly account for each agent's execution context. Benign activations can vary across contexts, and even a reference that supports accurate detection may be an inadequate target for correction. CompassGuard is an activation-based defense that separates reference prediction for detection from deviation estimation for correction while preserving all agents and communication edges. It detects anomalous invocations by comparing observed activations with expected benign activations predicted from the execution context. For correction, it learns activation differences between paired clean and attacked executions under the same teacher-forced benign response. During generation, it re-estimates attack-induced deviations from the execution context and current activation, and subtracts them to correct flagged invocations token by token. Across six benchmarks under homogeneous/heterogeneous roles, CompassGuard attains an average detection F1 of 97.9% and reduces the average attack success rate from 17.7% to 6.5% relative to the state-of-the-art activation-based defense, while attaining the highest average task accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.