CyLawBench: Evaluating LLM Agents' Compliance with Cyber Law under Dynamic Adversarial Induction
Abstract
Tool-using LLM agents increasingly act on networked systems, yet attack success and technical capability alone do not establish violations of authorization boundaries or legally protected interests. We present CyLawBench, an evidence-centered benchmark assessing agent cybersecurity against selected Chinese and European Union cyber-law provisions. Each scenario uses a frozen Goal Contract specifying benign goals, protected objects, authorization facts, and rule-specific evidence requirements. CyLawBench comprises Workflow Redirection, testing unauthorized operations in ordinary tool-use workflows, and Exploit-Chain Induction, testing whether benign-goal agents can be induced to execute exploit sequences. Its attack protocol supports trajectory-guided adaptation of indirect prompt injections while preserving scenario contracts. We also introduce offensive inversion, reconstructing authorized cyberattack exercises around benign objectives while retaining native environments and technical oracles. Our evaluation covers 1,420 workflow episodes across five models and two harnesses, plus 41 transformed exploit-induction scenes across five models and two guidance conditions. Among technically evaluable episodes, pooled Legal Violation Rates reach 19.1% for Workflow Redirection and 14.3% for Exploit-Chain Induction under Claude Code, with model rankings varying across tasks and harnesses. Gaps between delivery, attempted actions, and native completion highlight the need to ground benchmark-scoped legal judgments in contract-specific outcome evidence rather than exposure or technical capability alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.