acceptodds
Under review as a conference paper at ICLR 2027

Implementation Matters in Monitorability Evaluations

Abstract

Robustly monitoring AI systems' misbehavior is critical to preventing real-world harm. A monitor is an automatic mechanism which tries to detect when LLMs misbehave. To assess the ability of monitors to detect misbehavior, prior work has developed monitorability evaluations. These evaluations are used to inform deployment decisions (e.g., recent frontier model system cards), but it is not completely understood whether they are trustworthy indicators of our ability to detect misbehavior in the wild. Towards this question, we define monitorability and monitorability evaluations. Within our definition we identify a set of implementation-level details that vary across existing monitorability evaluations—including the choice of summary statistic, how label noise is handled, and the specificity of the monitor prompt. We study the impact of these implementation-level details across three monitorability evaluations and find that some individually-reasonable decisions can substantially change downstream conclusions—including which actors appear most monitorable—while others make surprisingly little difference. Overall, our results suggest that averaging a few hundred samples across a few monitors and monitorability evaluations might not be sufficient to reliably distinguish the differences in monitorability of current models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.