acceptodds
Under review as a conference paper at ICLR 2027

Self-Attribution Bias Makes AI Monitors Go Easy on Themselves

Abstract

Agentic systems increasingly rely on language models to monitor their own behavior. For example, coding agents may review generated code before approving a pull request or assess the safety of tool-use actions before continuing. Most static monitor evaluations present a fixed action to the model in a new context. In deployment, however, a monitor may instead evaluate an action that is already part of its own conversational trajectory. We show that these two ways of framing an action can produce different judgments: models often evaluate the same action more favorably when it appears in their own assistant trajectory compared to evaluating the same action in a new context. We call this effect self-attribution bias. Across agentic coding and computer use, monitors rate high-risk or low-correctness actions more favorably when they are self-attributed. For example, in SWE-bench code review, AUROC for separating passing from failing patches drops from 0.99 in a new context to 0.92 when the monitor rates its own patches in a previous assistant turn, because self-attribution inflates failing-patch ratings much more than passing-patch ratings. In contrast, explicitly stating that the action comes from the same model does not induce the same effect. This bias is much weaker when the evaluated actions come from other models, so static evaluations on fixed datasets can make monitors appear more reliable than they actually are in deployment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.