acceptodds
Under review as a conference paper at ICLR 2027

ActivationSMC: Sequential Monte Carlo for Optimizing Activation Steering

Abstract

Activation steering controls pretrained language models and beyond at inference by modifying internal activations, but its effectiveness hinges on jointly deciding where and how strongly to intervene. This creates a hard optimization over a mixed discrete-continuous space with coupled layer choices and often non-differentiable rewards, poorly handled by fixed grids or greedy decoupled search. We introduce \bf ActivationSMC, a derivative-free black-box framework based on Sequential Monte Carlo that maintains a population of full steering configurations. It guides the population toward high-reward regions via reweighting, resampling, and stochastic mutation that preserves diverse modes, with a multi-fidelity schedule that uses small subsets for cheap early exploration and larger subsets for precise refinement. Without updating model weights, ActivationSMC configures off-the-shelf activation steering operators such as COLD-Steer, AUSteer, CAA, and Spherical Steering. Across TruthfulQA, R-Bench, AMBER, as well as text-to-image generation using Stable Diffusion, ActivationSMC consistently outperforms native tuning. It improves collapsed baselines by up to 28.9 percentage points, raises compositional text-to-image alignment by 20–30 percentage points across four steering operators, and substantially widens its lead as joint multi-layer optimization exposes the breakdown of decoupled heuristics.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.