acceptodds
Under review as a conference paper at ICLR 2027

Cross-Model In-Context Neurofeedback Reveals (Limited) Activation Control in LLMs

Abstract

Probes on activation vectors are widely used to detect misbehavior in language models, but their success relies on the assumption that models cannot control what their activations encode. However, humans can learn, through neurofeedback, to control the firing of even single neurons; therefore, recent studies investigated whether this ability extends to language models, reaching conflicting conclusions. Notably, no experimental design so far disentangles competence at the task from genuine access to one's own activations. We fill this gap with in-context neurofeedback sessions, in which a "subject" model is asked to maximize (or minimize) a score derived from its activations, which we show back to it at every turn. In our first experiment, the subject generates the sentences to be scored. However, to disentangle competence at the task from actual control, we also let a second "writer" model generate them, while still deriving the score from the activations of the subject. We test ten open-weight models from five families, with sizes from 9B to 32B. In the first experiment, some models move their own scores considerably (even modulo controls), but no more than a writer of comparable competence moves them. Is this because succeeding at the task requires no privileged access, or because the subject controls its activations while processing the sentences, whoever wrote them? To answer this, we fix the sentence, prefilled as the subject's reply at every turn, and ask the subject to increase or decrease the score. Strikingly, three out of ten models show small but consistent control effects. We further ask whether models can control their activations to the point of fooling an activation monitor. To this aim, we make models repeat sentences a monitor would flag, such as a deceptive answer for a deception monitor, and ask them to raise or lower the monitor's score. No model learns to lower it, but, notably, one model can raise it. In-context control of a model's own activations might not yet be a threat to activation monitors, but our results show it exists in a few models, and it needs to be tracked as models evolve. Code and data will be released upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.