Activation Steering by Path Law Matching
Abstract
We develop a theoretical framework for activation steering based on matching activation path laws. Distinct path laws can share identical activation marginals at every layer, so marginal alignment alone cannot identify the intended path law. Using Kolmogorov's extension theorem, we characterize this law through consistent finite dimensional joint distributions. For finite depth networks, we construct the Kolmogorov loss to match joint activation distributions over contiguous groups of layers. Under the stated kernel conditions, vanishing population loss is equivalent to weak convergence toward the reference path law. With shared inputs and additional regularity assumptions, sufficiently accurate matching of actual intervened trajectories improves output agreement over exact marginal matching that retains positive functional error. We also establish conditions under which the difference-in-means vector is the unique optimal additive map. For practical fitting, we learn steering maps from cached activations without additional forward or backward passes through the pretrained model. Experiments on jailbreak and truthfulness across language models with 7B to 31B parameters demonstrate the framework's utility across several steering map families.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.