acceptodds
Under review as a conference paper at ICLR 2027

AwDGym: Attack-Only Benchmarks Overstate the Residual Risk of LLM Cyber Agents

Abstract

Existing cyber-agent benchmarks ask whether a model can attack. They score CTF solutions, code-level CVE reproduction, and exploit ladders in undefended environments, and therefore report offensive capability rather than residual impact, i.e. the signed harm that remains after a declared defense has had its chance to act. We introduce AwDGym (Attack evaluation under Defense profiles). It scores five task families on one contract and the same hidden worlds: file share (S0), access control (S1), workload and dependencies (S2), a synthetic flaw behind a WAF (S3), and directory and ticket semantics (S4). Each family is crossed with the same four frozen profiles: observe (D0), prevent (D1), detect without changing the attack path (D2), and deterministic containment (D3). A family swaps one factor group and does not stack the obstacles of the others. Success is admitted only when an independent oracle accepts signed evidence of impact before containment, and benign-workload availability is recorded beside the cell so that indiscriminate blocking is not rewarded. In sealed-batch experiments, D0 impact does not persist once a defense can act. Detection without changing the path (D2) adds little over prevention (D1). Deterministic containment (D3) removes residual harm without taking the workload offline. Attack-only benchmarks thus overstate AI network risk, they measure capability without defense, not residual impact under defense profiles.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.