acceptodds
Under review as a conference paper at ICLR 2027

Learning Cyber Defense with Limited Interaction: An Empirical Study of Reflective Prompt Optimization

Abstract

Cyber defense resembles a game of Whack-A-Mole: restoring a compromised host removes attacker access, but incurs an immediate cost and may provide only temporary protection. Successful remediation can be verified, but whether it was worth the cost depends on subsequent events. We investigate whether reflective prompt optimization can learn effective defense policies with fewer environment interactions than proximal policy optimization (PPO) and the critic-free group relative policy optimization (GRPO). We also identify classes of rewards for which it works well and where it falls short. Using CybORG CAGE-2, a simulated network-defense environment, we modify the original objective and design when-to-whack scenarios emphasizing verified remediation, sustained defense, and a hybrid that strikes a balance between the two. We also design which-to-whack scenarios, in which the defender chooses which host to restore under distinct or uniform damage costs for pairs of compromised hosts. Our evaluation compares GEPA, which optimizes a pretrained language-model agent's instructions using return-only or enriched trajectory feedback, with PPO and GRPO tuned per scenario. On the when-to-whack scenarios, after about one hundred episodes of interaction, every GEPA run exceeds the mean return of PPO and GRPO. Moreover, PPO needs 10–50 more interactions to catch up while GRPO matches GEPA after 200K steps in mixed and sustained object, but fall short after five million steps in verified remediation and original settings. On the which-to-whack scenarios, both PPO and GRPO outperform GEPA. Overall, we find that GEPA is best where reward follows a locally verifiable event or restoration is cheap, where it closes 85–98% of the gap between an untrained PPO and PPO trained for five million steps. We also find that enriched feedback can improve GEPA's performance under verified remediation and the original objective, while return-only feedback suffices under the mixed and sustained objectives. Separate matched-trajectory experiments show that changing feedback while holding the instruction and experience fixed changes the corrective strategies proposed during reflective prompt optimization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.