acceptodds
Under review as a conference paper at ICLR 2027

The Monitorability Geometry of Language Models

Abstract

Linear activation monitors, whether probes, subspace monitors, or a model's own self-report, read only a few directions of a high-dimensional hidden state, so something always falls in their blind spot. We give a label-free, closed-form way to measure how much control over a model's behavior survives outside any rank- monitor, from two Jacobians taken at one hidden state, one for what the model does and one for what it says about itself. From their geometry we bound what a hidden intervention can achieve, find the best monitoring subspace, and connect visibility to detection. One geometry governs both sides. A monitor fixed before deployment captures of a model's behavioral leverage at rank and at rank (Llama, Qwen3.5, Gemma, and OLMo; 1B–70B), but rank is not the lever. Against an attacker who accepts visibility, widening from rank to lowers guaranteed evasion only from to . What visibility concedes, no rank buys back. The blind spot is concrete. A steer along the refusal direction a rank- activation-PCA monitor cannot see flips refusal on AdvBench across four families while staying invisible (), whereas the same steer in the monitor's view barely moves it. The cause is that the variance directions these reads occupy align poorly with what drives behavior. On Gemma-4-12B a rank- PCA read captures of the refusal direction, below the random null, against for the matched Jacobian subspace. Spending that visibility is also what catches an attacker. Once the steer is strong enough to flip refusal, a Fisher-based monitor detects it (AUROC –) where PCA stays blind (). Introspection follows the same geometry in sharper form. – of the behavior-driving direction lies outside the self-report subspace, so a self-report can be forged, moved to say one thing while behavior stays fixed. A single overlap between the report and behavior directions predicts how faithful that forgery can be, which we confirm causally by fine-tuning models to move it. The same account predicts when a model's report that it is being evaluated comes apart from how it acts. Together these results give a concrete calculus for choosing a monitor's width, threshold, placement, and subspace, and for knowing what it must miss.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.