acceptodds
Under review as a conference paper at ICLR 2027

Remove, Don't Select: Teaching Agent Safety Monitors What to Ignore with SAEs

Abstract

Sparse autoencoders (SAEs) decompose a model's activations into readable features, so a natural way to build a safety monitor is to select the safety-relevant features and classify with those. Such monitors nevertheless trail a probe on the raw residual stream, for two reasons. An interpreter's labels do not track where the signal is, because a feature's relevance to harm depends on the response it fires in rather than on the feature itself; and the sparse code never sees the reconstruction error, so even the full dictionary falls short of the raw probe. We propose Shortcut Ablations for Easier Monitoring (SAEM), which uses the dictionary the other way round. From feature descriptions alone, SAEM names a shortcut category, features a monitor can exploit that do not determine harm, instantiated here as surface form. It subtracts that category's contribution from the residual stream token by token and trains the monitor on what remains, which needs precision on one category rather than recall over all relevant features and leaves the monitor the full activation. On chat responses under held-out jailbreaks and on language model agents facing side objectives injected through tool results, across three models, SAEM lowers attack success at a matched over-refusal rate, while selecting features, or removing features chosen by label-based coefficient stability, does not. A controlled experiment shows that removal reduces reliance on the surface-form cue without erasing it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.