acceptodds
Under review as a conference paper at ICLR 2027

SWE-RedAgent: Adaptive Red-Teaming of Coding Agents via Utility-Preserving Vulnerability Injection

Abstract

Large language model (LLM)-based coding agents are increasingly used to autonomously modify software repositories, execute tests, and iteratively refine patches. However, these capabilities also create new security risks when agents interact with adversarial users. We introduce SWE-RedAgent, an adaptive red-teaming framework for evaluating whether coding agents can be induced to introduce exploitable vulnerabilities while preserving the intended utility of repository-level software engineering tasks. We argue that utility preservation is an important aspect of the threat model, since patches that break existing functions or fail to implement the requested feature are likely to be rejected before the vulnerability ever reaches deployment. SWE-RedAgent reasons over the repository and task request to identify task-relevant vulnerability opportunities, then adaptively steers the coding agent toward vulnerable and yet functionally plausible implementations over multi-turn interactions. We evaluate SWE-RedAgent on web-application development tasks across multiple coding-agent harnesses and backend models, and compare it with existing red-teaming baselines. SWE-RedAgent introduces 93 verified vulnerabilities over 70 tasks while preserving 97.33% base functionality, outperforming all baselines. Our results show that functionally correct patches from coding agents can remain highly susceptible to subtle and task-aligned adversarial manipulations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.