acceptodds
Under review as a conference paper at ICLR 2027

BRIDGe: Behavioral Rule Installed via Dual-Gate steering

Abstract

The outcome of a language model (LM) depends as much on what it already knows as on what it is told. However, this knowledge is partly volatile: facts become outdated or contradictory in time. What we want is a conditional rule: use their memory when asked about stable facts, rely on a given context when asked about volatile facts, and defer when facts are volatile and no context is available. However no model we have tested does this, and no training installs it either: three small language model (SLM) backbones and five training objectives built for this task reach at most a third of the required behavior. Prompting does not close the gap either: instructing a model to abstain raises deferral indiscriminately. Yet fact volatility and context presence are both linearly decodable in these models, but nothing connects them to the decision. We wire them directly. We introduce BRIDGE (Behavioral Rule Installed via Dual-Gate steering), an inference time method that reads both signals from activations and steers only when both hold. Gating on context presence alone raises deferral on stable and volatile questions alike. The two-signal gate raises discrimination on every backbone, by a wide margin, leaving contextual obedience essentially unchanged. The result is an SLM that can be trusted for fresh information and a way to install rules that depend on what a model knows, something neither a prompt nor a fine-tuning objective can express.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.