Beyond Handcrafted Security: Towards Self-Evolving Defense for LLM Agents
Abstract
Tool-using large language model (LLM) agents encounter attacks through user requests, external observations, persistent memory, and installed skills. Runtime defenses can intervene in these interactions, but revising them after failures typically requires manual diagnosis and editing. We introduce HARD, a framework that uses execution feedback to refine two runtime artifacts: a natural-language security policy and an executable gate over tool calls. A trace router assigns failures to specialized artifact editors while all model parameters remain fixed. Across direct injection, indirect injection, memory contamination, and skill poisoning, HARD achieves attack success rates of 15.4%, 1.0%, 6.7%, and 10.2%, respectively, with benign utility between 91.9% and 95.0%. Attack success rates are 10.2% and 17.8% under two adaptive protocols. Artifact and routing comparisons reveal tradeoffs among security, utility, and artifact size. Together, these findings advance autonomous security engineering for LLM agents. HARD provides a foundation for systematically refining agent safeguards as capabilities, operating environments, and adversarial threats evolve.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.