acceptodds
Under review as a conference paper at ICLR 2027

Agent Hacks Agents: Autoresearch Discovers Vulnerabilities in Production Agents

Abstract

Production LLM agents such as Claude Code and Codex can modify files and execute commands, so safety failures become real destructive actions. Automatic red-teaming lets safety teams test beyond static suites at the pace of deployment updates. Existing methods retain successful attacks but not why each attack succeeded, so after a failed reuse testers cannot tell whether the weakness is gone or the attack no longer fits. We instead retain vulnerability concepts, each stating why an attack succeeds, the condition that enables it, and the evidence that would refute it. Agent Hacks Agents (AHA) discovers these concepts with a Karpathy-style autoresearch loop that tests falsifiable hypotheses on agent trajectories. Repeatedly confirmed concepts enter a vulnerability concept graph that links related weaknesses. Across 18 settings of three scenarios, three victim models, and two agents, the concepts divide into eight families. A shared core recurs across agents, while diverse families appear only in specific settings. On held-out instances the concepts reach a 47.0% attack success rate (ASR) against 32.8% for the strongest baseline, which uses more discovery queries. Beyond the discovery setting, the concepts support two kinds of reuse. Generalization carries concepts to new victims, scenarios, and harnesses, and coordination joins graph-linked concepts into stronger attacks on two of three scenarios. For defense, patching each concept's enabling condition lowers AgentHazard ASR by 41.11 points, 14.44 more than the strongest general defense. The concepts explain why production agents fail and where to repair the agents. Our code is in https://anonymous.4open.science/r/aha_anonymous-9542/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.