acceptodds
Under review as a conference paper at ICLR 2027

Target-Distribution Matching for Activation Steering: Paired Causal Residual Transport

Abstract

Activation steering changes a frozen language model’s behavior, but behavioral scores and hidden-state distances do not specify which output distribution it should produce. For instruction-defined control, the model supplies this target: its next-token distribution under the target instruction, given the same content and prefix. We introduce Paired Causal Residual Transport (PCRT), which predicts the same-content residual left by a fixed editor and adds it after final normalization, before the vocabulary head. The target instruction supplies training supervision and is absent at deployment. Within the learned residual subspace, the conditional mean minimizes squared output-head error, which bounds probability error before numerical rounding. On Gemma 2 2B, PCRT reduces 16-token sequence Kullback–Leibler (KL) divergence by 35.7% to 54.4% for Activation Transport (AcT) and Linear End-to-end Activation Steering (LinEAS), outperforming matched different-content residuals on all 64 test contents. AcT repair reduces 32-token sequence KL by 69.6% to 74.2% on Qwen 2.5, Qwen 3, and Llama 3.1. On five public instruction-following categories, direct correction lowers macro sequence KL by 22.2% and raises strict compliance by 44.1 percentage points. A fixed editor’s remaining distributional error can therefore provide a learnable correction target, without retraining the editor.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.