acceptodds
Under review as a conference paper at ICLR 2027

Auditing the Auditor: Ground-Truth Controls for Natural-Language Explanations of LLM Activations

Abstract

Natural Language Autoencoders (NLAs) generate natural-language explanations of language-model activations and have been proposed as an auditing tool for unverbalized properties such as evaluation awareness. Because the explainer is itself a language model trained for reconstruction rather than faithfulness, a specific-sounding explanation is not by itself evidence about the activation, and the method's authors do not establish calibration of their awareness measurements. We contribute a ground-truth calibration harness for selected released NLA checkpoints (Gemma-3-12B/27B; earlier, Qwen2.5-7B) that scores explanations against facts or conditions we specify, or against a word the model later reveals, under preregistered controls: a wrong-activation (shuffled) control, a visible-text baseline, a blinded judge, and supervised readability probes. Exploratory experiments show that apparent successes on famous facts are indistinguishable from verbalizer prior-completion under a truthful-instruction control and do not recur when the facts are invented (a thin positive control: one cleanly reported fact per scale). In a preregistered hidden-condition experiment, an assigned evaluation-vs-ordinary mode was linearly decoded at both scales, although placebo controls do not rule out token echo or instruction residue; keyword screening nominated 1/31 (12B) and 1/32 (27B) items, and neither met an amended blinded-human reporting criterion (two independent readers; nominated cells not resampled). In a preregistered prospective-reveal experiment on Gemma-3-27B, no early-position free-form NLA explanation contained a candidate word (0/96), while the target model, choosing between the two allowed words from the visible conversation, predicted the later reveal in 61/96 cases (exact McNemar ); a supervised probe did not establish that the reveal was recoverable at that position. Under the wrong-activation control, explanations contained the substituted item's word far more often than the item's own (147 vs. 5 of 576 texts, by literal occurrence), almost entirely immediately before the reveal. We present the control stack as a reusable harness so that future explainers can be evaluated against the same standard.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.