ProtoArena: Can AI Agents Discover and Exploit Real-World Cross-Component Protocol Vulnerabilities?
Abstract
Large language model (LLM) agents are increasingly used for vulnerability analysis and automated security testing. Existing cybersecurity benchmarks primarily evaluate capture-the-flag challenges, source-assisted vulnerability reproduction, and exploitation of vulnerabilities within software systems. However, many protocol vulnerabilities arise from inconsistencies in message interpretation, timing, state, or trust assumptions across multiple components. Existing benchmarks provide limited coverage of whether agents can infer and exploit such interactions through black-box observation. We introduce ProtoArena, a black-box, multi-component, dynamic benchmark comprising 80 cross-component vulnerability-exploitation environments across 11 protocol and application domains and seven attack categories. Its tasks are primarily informed by published protocol-attack mechanisms. Agents interact with running systems through scoped interfaces to achieve observable security effects without access to target source code. Three information levels assess guided reproduction, exploitation with mechanism and topology context, and effect-directed black-box exploration. A hybrid verifier combines deterministic checks with LLM review of independently collected traffic and logs. We evaluate five LLMs on ProtoArena. Claude Opus 4.8 achieves the highest observed success rates, reaching 48.8% under guided reproduction and 12.5% under effect-directed black-box exploration. In a separate DNSBomb optimization experiment, it achieves a peak downstream-response to average query-input byte-rate ratio exceeding . These results demonstrate verifiable exploitation in a subset of the evaluated environments, while revealing substantial limitations under reduced guidance and variation across protocol domains and attack categories. All experiments use controlled, explicitly authorized targets and publicly documented attack mechanisms; they neither test production systems nor disclose previously unknown vulnerabilities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.