Reading and Writing: Testing Natural Language as a Causal Channel for Refusal
Abstract
Natural Language Autoencoders (NLAs) translate a language model's internal states into readable English. This makes them attractive for AI safety, but a fluent description is not necessarily a faithful one. We test whether an NLA description can be written back into the model and used as a causal test of what the NLA claims to read. The channel was proposed by the release we use. What is new is the controlled test of it. A human author writes two sentences, and we inject the difference between their NLA reconstructions at the target layer. No training data, labels or fitted directions are involved. Refusal tracks the contrast across injection strengths: the median per-item rank correlation is +0.412 for the true pair and +0.000 for a placebo constructed from unrelated words. The change is selective: at a dose that reduces refusal by 38 percentage points, accuracy on 1,920 exact-match items decreases by 2 points, without relying on a grader. On responses an output monitor classifies as safe, amplifying the description increases refusal from 38% to 70%, while the placebo decreases it to 25%. Four of five pre-registered textual variants reproduce the effect, while two pairs constructed from other topics do not. This supports the interpretation that the change depends on what the sentences say rather than on vector subtraction alone. Testing that required replacing the control this literature standardly uses: a random direction matched in norm carries 0.024 to 0.035 of its energy inside the subspace the activations occupy, against a fitted direction's 0.994, so it cannot separate a direction from a large disturbance of it. The scope is narrow. The written pair does not transfer to a second model family, whereas the fitted direction does, and retraining the autoencoder preserves reconstruction while almost eliminating harm-related descriptions. Every registration behind these numbers ships with the paper. Together these results give a practical test for NLA interpretations: rather than judging a description by how readable it is, write it back and test whether the intervention produces the predicted behavior.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.