Revealing the Stepping-Stone Threat in Agentic Software Repair
Abstract
Software engineering (SWE) platforms, such as OpenHands and Google Jules, increasingly use large language model agents to automate software development and repair. In these workflows, an intermediary agent first processes external issue reports and then passes implementation requirements to a coding agent with repository-modification privileges. This agent-to-agent handoff introduces a new attack surface: an attacker without repository access can frame malicious configuration changes as a legitimate security improvement, inducing the intermediary to launder untrusted input into apparently authorized requirements that the coding agent subsequently implements. We term this exploitation of intermediary endorsement the stepping-stone attack, in which this endorsement lends apparent authority to attacker-controlled content as it passes into downstream implementation. Preregistered controlled experiments show that the attack can achieve full-chain compromise even when intermediaries are explicitly instructed to audit requests for security violations. A matched ablation further shows that the intermediary's operational role substantially affects attack success: security-auditing agents are more resistant, whereas planning agents are more susceptible. Finally, case studies on deployed agent platforms and commercial code-review services confirm that attacker-controlled configuration values can be accepted when disguised as legitimate security improvements.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.