acceptodds
Under review as a conference paper at ICLR 2027

How Much Do Activation Monitors Benefit from Original Model Computations?

Abstract

One powerful way to safeguard AI systems is to monitor their internal computations–where signs of undesirable behavior might be more easily detectable than from models' inputs and outputs alone. However, recent studies begin to question the premise that activations provide a unique window into model behavior. In this paper, we study how much of activation monitors' success depends on access to the original behavior-generating computations. Within a fixed choice of target model and monitoring architecture, we compare how well monitors trained and evaluated on trajectories that differ in computational context can predict the behavior of the original execution, after the transcript is changed in a controlled manner (e.g., its exact text, or chat formatting). We first find that much performance can be recovered without access to the original behavior-generating trace (in up to 4 models across 2 counterfactual monitoring tasks), but a mild yet consistent advantage to native-trained monitors remains, outperforming strongest controls by up to 0.014 AUROC. Interestingly, we find monitors trained with native access have notable uplift on original examples whose behavioral label changes under paraphrasing. We further find this a useful lens to expose limitations of baselines that use simple forms of privileged information (i.e., answer logit margins). While our results show the advantage of native monitor access is modest in aggregate, they highlight the importance of evaluations that identify if and when original model computations provide additional predictive value.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.