Activation Steering Transfer to Agents: One Gain Ratio Does Not Identify Potency and Efficacy
Abstract
Activation Steering Transfer to Agents: One Gain Ratio Does Not Identify Potency and Efficacy Additive activation steering is calibrated in single-turn chat and deployed inside agent scaffolds. The quantity usually reported for that move is a gain: a ratio of steered effects, T=Delta_agent/Delta_chat. It is a summary of the same kind as a ratio of Braun's steerability slopes (Braun, 2026), though not an identity with one: the two coincide only under the proportional linear response this paper rejects. It is structurally a member of the relative-steerability family of Tan et al. (2024). We show on eight family x arm dose-response cells over six models, swept in both deployment contexts, that this ratio does not identify potency and efficacy: reconstructing the published estimator in both of its forms on our own grids, its realized range contains 1 in five of five scorable cells, it moves with dose in four of five, and two cells with opposite potency shifts, both CI-clean on the primary grid, return gain intervals that overlap at a width under 0.08. Every scorable cell in the frozen five-cell gain audit is an amplifier at one dose and an attenuator at another, so an amplify/attenuate taxonomy reports the dose it was read at. We replace the gain with a location: dEC50 = EC_50^agent - EC_50^chat, the cross-context difference in a curve location whose context-dependent horizontal reparametrization Taimeskhanov et al. (2026) derive analytically, in a simplified model, for one context at a time. It is a curve location rather than a ratio of effects, signed in both directions on this roster's primary grid across four distinct model families (+1.013[+0.777,+1.273] against -12.368, -10.855, -5.497 and -886.066 elsewhere), and beats a vertical rescaling at equal complexity in all five cells of the frozen horizontality audit (four under the registered alpha^*-trim). Because separating a mechanism from an artifact is the question at issue, we report the discipline at the same volume as the result: one cell is quarantined loudly, because its agent response inverts past a coherence collapse and the fit's orientation flips, manufacturing the roster's largest positive, which we delete. Three deflationary accounts (residual-norm rescaling, baseline alignment, dose transduction) are measured and rejected on sign pattern and magnitude, including wrong-sign predictions on the roster's positive anchor. Our pre-registered forecaster of dEC50 was refuted out of sample and published as refuted, our registered salvage claim (a location estimable where a gain ratio dies against a ceiling) produced no such cell and is reported unanswered, and a census of our own pre-registration register reports the registered branches no code in this repository could have emitted. The practical consequence is a measurement instruction rather than a theorem: a transfer conclusion read at one strength does not identify what changed, because a displacement and a gain are not distinguishable from a single operating point.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.