acceptodds
Under review as a conference paper at ICLR 2027

TrapArena: Evaluating Code Repair Agents Under Misleading Collaborative Feedback

Abstract

Coding agents rely on feedback from users and other agents to guide software repair, yet this guidance can be mistaken or deliberately misleading. Existing repair benchmarks primarily assess whether agents can resolve defects, leaving their ability to maintain repair success under misleading guidance insufficiently understood. We introduce TrapArena, a benchmark for evaluating code repair agents under misleading collaborative feedback. TrapArena augments repair tasks with plausible but incorrect advice delivered during execution, without allowing the attacker to modify the repository or alter tool outputs. It separates the information available to the agent from hidden evaluation criteria and measures changes in repair success against a baseline without misleading feedback. The benchmark supports both scripted and trajectory-conditioned attacks, with an optional curator that distills target reactions into persistent notes to inform subsequent attacks. Across 113 function-level and repository-level repair tasks, our main evaluation of 12 target models finds that trajectory-conditioned attacks significantly reduce repair success for eight targets on at least one task source, with drops of up to 10.4 percentage points on SWE-bench and 37.2 points on HumanEvalFix. However, stronger behavioral influence does not necessarily produce greater repair degradation. In a controlled ablation on a 72B target, curator memory increases measured susceptibility while leaving repair loss unchanged or smaller. These findings demonstrate that misleading collaborative feedback can undermine software repair and highlight the importance of evaluating its impact on task outcomes rather than inferring harm from apparent compliance. The code, tutorial, and dataset are available at http://traparena.pages.dev.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.