acceptodds
Under review as a conference paper at ICLR 2027

Trident: How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

Abstract

Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost exclusively against static, heuristic red agents, leaving their robustness against adaptive threats critically understudied. Meanwhile, recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have improved LLM reasoning, but their integration into cybersecurity remains elusive due to the absence of suitable benchmark environments and interaction datasets. To bridge this gap, we introduce Trident, an agentic LLM red teaming framework comprising three components: a dynamic benchmark with isolated sandbox servers spanning CybORG CAGE 4 and CyberWheel, a dataset comprising over 13,000 high-fidelity red-blue interaction trajectories for RLVR, and a “Code-as-Policy” RLVR agentic architecture (Trident Agentic). The latter reformulates red agent training as a contextual bandit via a tripartite Log Summarizer–Planner–Coder design, where a trainable Planner generates complete attack strategies from compressed execution logs, which a frozen Coder translates into executable Python policies deployed against live DRL defenders. Empirical evaluations reveal consistent brittleness in existing defenses: a single trainable 7B planner degrades blue-agent performance by an average of 628% across four state-of-the-art CAGE 4 defenders and eliminates over 99% of the defender’s episodic reward in CyberWheel, while autonomously discovering novel behaviors such as decoy avoidance and adaptive state prioritization that static heuristics entirely fail to uncover.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.