acceptodds
Under review as a conference paper at ICLR 2027

Language models recognize dropout and Gaussian noise applied to their activations

Abstract

We provide evidence that language models can detect, localize and, to a certain degree, verbalize the difference between a-semantic perturbations applied to their activations. More precisely, we either (a) apply *dropout* or (b) add *Gaussian noise* to them, at a target sentence. We then ask a multiple-choice question such as *Which of the previous sentences was perturbed?* or *Which perturbation was applied?*. We test eight models from the Gemma, Llama, Mistral, Olmo, Qwen and Seed families, with sizes between 8B and 36B, all of which can easily detect and localize the perturbations, often with perfect accuracy. In studying the mechanisms behind this capacity, we find that models develop it already at pretraining, and they do not simply track the log-likelihood of the perturbed sentences. To investigate the mechanisms, we find that Olmo models develop this capacity during pre-training, and we rule out the hypothesis that perturbed sentences simply have lower log-likelihood. Notably, the zero-shot accuracy of Qwen3-32B and Qwen3.8-27B in identifying *which* perturbation was applied correlates with the perturbation strength. While this task proves hard for the remaining models, most of them can learn, when taught in context, to distinguish between dropout and Gaussian noise. Interestingly, Qwen3-32B and Qwen3.8-27B learn worse when supervised with wrong labels, again suggesting a prior towards the correct ones. Because dropout is used as a training-regularization technique, while Gaussian noise is sometimes added during inference, we discuss the possibility of a data-agnostic *training awareness* signal and the implications for AI safety.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.