acceptodds
Under review as a conference paper at ICLR 2027

I-State: A Self-Regulating Controller for Transformer Language Models

Abstract

A language model's resistance to manipulation can be measured, but it cannot be certified. We built a controller whose off-state we could prove exact, then found it can never reach that state — and closed the gap. I-State attaches to a frozen model a typed state on three timescales: a protected core the conversation cannot edit, a mood that decays at a fixed rate, and a memory of sustained pressure. It reads the model's own activations and sets its own strength — no dial, no human in the loop. Because one gate multiplies every trainable term — a FiLM modulation, a low-rank adapter and a refusal direction — a calm model is the base model on identical tokens: every measurement is a per-example counterfactual, not an estimate. On a benchmark we built for multi-turn persuasion against non-harmful deployment policies, the model holds its policy .25 to .52 more often than the frozen base across two model families from 1.5B to 14B, over bases holding .11 to .36. Removing the dynamics alone costs a third to two-thirds of that gain across three retrainings. The affect map has trivial kernel, so no readable input drives the gate to zero: uncorrected it idles between attacks, at 2.4 MMLU points. Re-centring the origin returns benign behaviour to the base model's own rate at unchanged violation, without retraining, and brings the deployed cost to 0.1 MMLU points — against 2.8 for the best always-on constant. An off-state a human sets is a switch; only one the dynamics reach is a policy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.