acceptodds
Under review as a conference paper at ICLR 2027

Comparable Outcome, Different Execution: How Should We Measure Evaluation Awareness in Coding Agents?

Abstract

Coding agents operate through long trajectories in which they inspect repositories, use tools, modify code, and encounter information that may reveal how they are being evaluated. We introduce SWE-EvalAware, a benchmark of 910 controlled software-engineering task variants spanning 13 controlled settings per SWE-Atlas base task. Across four coding-agent stacks and 9,444 usable trajectories, we hold the underlying software task fixed while varying evaluation cues, consequences, resource constraints, and message placement, and separately measure information exposure, expressed evaluation-related reasoning, execution behavior, and task outcome. Evaluation awareness is heterogeneous rather than binary. Our trajectory-level taxonomy identifies seven non-mutually-exclusive categories, with provenance and evaluation-apparatus awareness substantially more common than other forms. Awareness also emerges without planted evaluation cues, including when agents reconstruct task provenance or discover benchmark-specific information during execution. Repeated sampling increases observed awareness from 50.6% at one attempt to 63.1% at three, but this increase closely tracks the gain in task success and masks large differences across awareness categories. Provenance-aware runs are associated with substantially more external lookup and higher task success, while integrity/reference awareness is also associated with higher success. We further observe direct retrieval and application of public benchmark reference solutions. Finally, consequence and resource controls show that large behavioral changes can occur without a corresponding increase in broader evaluation awareness. Our results show that evaluation awareness in long-horizon coding agents cannot be inferred from task success or behavioral change alone; reliable evaluation must separately measure what agents encounter, what they recognize, and how they act.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.