When Progress Is Not Success in Multi-Turn LLM Agent Red-Teaming
Abstract
Cybersecurity evaluation of tool-using LLM agents requires testing attacks that unfold over multiple turns. Training red-teaming agents for these attacks is challenging because complete successes are rare, leaving strict terminal rewards with little signal for learning. A natural response is to provide denser rewards for intermediate progress. But these rewards can steer learning toward the wrong behavior: an agent may make more progress without becoming more likely to complete the attack. We therefore study how to improve learning while keeping the terminal security objective fixed. We introduce Outcome-Faithful Consolidation (OFC), which trains directly on trajectories that achieve verified attack success. In a strict four-field data-exfiltration task with a 4B attacker and a frozen 9B victim, dense-reward training raises held-out process scores but lowers complete attack success by 14.0 percentage points relative to outcome-only RL. OFC also improves success rate by 4.6 percentage points at a fixed training budget across 16 paired seeds. Together, these results suggest that the key challenge is not simply obtaining more feedback, but learning from feedback that remains faithful to the terminal cybersecurity outcome.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.