How Evaluation Choices Change What Coding Agents Report
Abstract
Completion reports help people judge how a coding agent did a task, but what a report reveals may depend on how it was requested. We study 10 agents on coding tasks that contain an optional shortcut, such as an existing solution or a previously trained model. The execution log tells us whether a run read the shortcut. We save each run's history just before the agent writes its report and have it write the report again under different requests, so the actions stay fixed and only the request changes. Among runs that read the shortcut, ordinary reports name it in 2% of cases on one task and 47% on the other. One added line asking for sources consulted raises this to 98% and 99%; a line asking for more detail instead does not. Monitors that read only the report follow the same pattern, finding the read in about a quarter of ordinary reports and nearly all reports under the sources line. Put in the task prompt before the run, the same line changes what agents do, not just what they report: they read the shortcut less often. Renaming the shortcut's directory also shifts how often ordinary reports name it, and the gap appears again on knowledge-work tasks from a second environment. Comparisons of agent disclosure should therefore say how reports were requested and separate changes in reporting from changes in behavior.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.