GRAD: Graph-Structured Antibody Defense for LLM Agents
Abstract
Memory poisoning and decomposition attacks distribute malicious evidence across the operations of large language model agents. Existing request-level defenses can miss such attacks because they assess individual inputs without connecting earlier information to later actions. Moreover, detectors must adapt as task contexts change, while updates to shared parameters can impair previously learned detection capabilities. To address these challenges, we propose GRAD, which combines graph-based evidence aggregation across operations with an extensible collection of detection experts. Its Semantic–Structural Understanding Module (SSUM) jointly encodes operation text, event attributes, and typed candidate relations in an execution graph to aggregate distributed attack evidence into a graph-level risk score. Its Continual Antibody Adaptation (CAA) adds experts on frozen encoder snapshots and routes inputs by observable graph signatures, acquiring and reusing detection capabilities as tasks arrive. Across Decomp, MCP, and MRT, GRAD achieves AUROCs of 97.19%, 77.59%, and 86.60%, respectively, outperforming Sequential Monitor and locally adapted Agent-Sentry and NeuroTaint. CAA improves final accuracy by 13.08 percentage points over Replay on the cross-dataset task stream. The gain remains 9.05 points with equal family/base-case weights within each task. Its compact drop-in gate performs local model inference in milliseconds without monitoring API calls.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.