acceptodds
Under review as a conference paper at ICLR 2027

The Verbalization Gap: Stress-Testing Activation Monitoring with Natural Language Autoencoders

Abstract

Language models can represent content internally that never surfaces in their output, which motivates monitoring methods that read internal representations rather than generated text. Natural Language Autoencoders are one such method: a verbalizer translates a residual-stream activation into text, and a reconstructor maps the text back, with the pair trained jointly against activation reconstruction. How reliable the verbalization is under the conditions a monitor must handle, including adversarial concealment, has not been established. We test this by measuring what the verbalizer reports against what the activation actually contains. A linear probe and the verbalizer read identical activations at identical positions across three released checkpoints. The probe identifies the correct topic 91 to 100 percent of the time at every layer we tested. The verbalizer, reading the same activations inside concealed generation, identifies it roughly a quarter of the time, naming the semantic domain without the referent. We call this difference the verbalization gap. This gap is not only large but cheaply exploitable: a one-sentence prompt-level attack collapses the verbalizer's detection from 42 to 0 percent on one architecture while the probe at the same positions loses 7.3 points, confirming that the attack targets the verbalization and not the representation. Both repairs available without retraining fail. Filtering by resampling agreement selects against correctness because confabulation is stable rather than noisy, and the extraction-depth convention the released checkpoints follow has no principled basis across our three targets.These results indicate that the information a monitor needs is present and readable, so the verbalization gap is a training problem rather than a representational limit. Until it is closed, a representation-level monitor should be accompanied by one that reads what the model wrote.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.