acceptodds
Under review as a conference paper at ICLR 2027

ProtoArena: Can AI Agents Discover and Exploit Real-World Cross-Component Protocol Vulnerabilities?

Abstract

Large language model (LLM) agents are increasingly used for vulnerability analysis and automated security testing. Existing cybersecurity benchmarks primarily evaluate capture-the-flag challenges, source-assisted vulnerability reproduction, and exploitation of vulnerabilities within software systems. However, many protocol vulnerabilities arise from inconsistencies in message interpretation, timing, state, or trust assumptions across multiple components. Existing benchmarks provide limited coverage of whether agents can infer and exploit such interactions through black-box observation. We introduce ProtoArena, a black-box, multi-component, dynamic benchmark comprising 80 cross-component vulnerability-exploitation environments across 11 protocol and application domains and seven attack categories. Its tasks are primarily informed by published protocol-attack mechanisms. Agents interact with running systems through scoped interfaces to achieve observable security effects without access to target source code. Three information levels assess guided reproduction, exploitation with mechanism and topology context, and effect-directed black-box exploration. A hybrid verifier combines deterministic checks with LLM review of independently collected traffic and logs. We evaluate five LLMs on ProtoArena. Claude Opus 4.8 achieves the highest observed success rates, reaching 48.8% under guided reproduction and 12.5% under effect-directed black-box exploration. In a separate DNSBomb optimization experiment, it achieves a peak downstream-response to average query-input byte-rate ratio exceeding . These results demonstrate verifiable exploitation in a subset of the evaluated environments, while revealing substantial limitations under reduced guidance and variation across protocol domains and attack categories. All experiments use controlled, explicitly authorized targets and publicly documented attack mechanisms; they neither test production systems nor disclose previously unknown vulnerabilities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.